v13.0.3 upgrade from v12: old hiddify-singbox keeps port 2000 so hiddify-core crash-loops; haproxy 3.4 crashes on reload; rawhttp-HC/naive-h2 presets fail #5566
Description
Activity
Same upgrade path here (13.0.3 from 12.x, non-docker, Ubuntu 24.04). I can confirm §1, and I think I found the exact root cause — plus a second, independent cause of the same
address already in usesymptom.Root cause of the stale units: the cleanup guards use
[ -e ]on a symlinkAll three teardown blocks for legacy units are guarded the same way:
scripts/migrate_layout.sh:50(rename_singbox_unit) —[ -e /etc/systemd/system/hiddify-singbox.service ]services/hiddify-core/run.sh:6— same test, same unitservices/panel/install.sh:29—[ -e /etc/systemd/system/hiddify-panel-background-tasks.service ]
/etc/systemd/system/hiddify-*.serviceis a symlink into the versioned tree (e.g.-> /opt/hiddify-manager/hiddify-panel/hiddify-panel-background-tasks.service).[ -e ]dereferences the link, so once the same upgrade has already deleted the old directory the test is false, and the whole block — thesystemctl disable --nowand therm -f— is skipped. The guard fails in exactly the situation it was written for.One-line fix in all three places:
if [ -e /etc/systemd/system/<unit> ] || [ -L /etc/systemd/system/<unit> ]; then
That is also why this is easy to miss: since the
rm -fnever runs, the dangling link stays and systemd keeps the unit asLoadState=not-foundwithActiveState=active, sosystemctl statusandsystemctl list-unitsboth look healthy while a v12 process is still running.The
hiddify-panel-background-taskscase specificallyBesides the singbox process holding
127.0.0.1:2000, the v12 celery worker+beat pair from the deletedhiddify-panel/was still running after 29 days, executing pre-MySQL code against the new MySQL-backed panel. Symptoms:MySQLdb OperationalError (2013, 'Lost connection to MySQL server during query')inhiddify_panel.err.logevery few minutes- every user's
last_onlineandcurrent_usagefrozen at the upgrade timestamp — the panel reported 0 users online while ~780 established proxy connections and ~0.5 MiB/s were actually flowing
Killing the two PIDs and removing the dangling link fixed it. The in-process APScheduler then picked accounting back up on its own at roughly a one-minute interval, so no cron was needed — the folding-in works, it just never gets a chance while the old worker is still alive.
Second, independent cause of the same
address already in useWorth splitting into its own issue if you prefer — it is not upgrade-specific and should be reproducible on a clean install.
After clearing the orphan,
hiddify-corestill crash-looped, failing on a different port every restart (49578,49886,49138,48412, …). The generated inbound ports for bothxray.jsonandhiddify-core.jsonsit inside the kernel's default ephemeral range:net.ipv4.ip_local_port_range = 32768 60999 net.ipv4.ip_local_reserved_ports = # empty27 of my 31
hiddify-coreinbound ports are inside that range. So any outbound socket on the box can be handed one of those ports as its source port, andhiddify-corethen loses the race on startup.scripts/common/sysctl.conftunes a lot of TCP knobs but never touches either of these two.The actual squatter turned out to be
rpxy-l4: it leaksCLOSE-WAITsockets to its own target (127.0.0.1:<ephemeral> -> 127.0.0.1:901) — about 970 of them at the time.CLOSE-WAIThas no kernel timeout, so those ports stay taken until the process is restarted.Two separate things would help:
- Add the generated inbound ports to
net.ipv4.ip_local_reserved_portsinscripts/common/sysctl.conf, or allocate them below 32768. Note that reserving only blocks future automatic allocation, so on an upgrade a process that already leaked into the range still has to be restarted once. - The
rpxy-l4CLOSE-WAITleak looks like a socket-lifetime bug in its own right.
What I am running as a workaround:
# /etc/sysctl.d/99-hiddify-reserved-ports.conf net.ipv4.ip_local_reserved_ports = <generated inbound ports, e.g. 47900-49960>then restarting
rpxy-l4once to drop the already-leaked sockets before startinghiddify-core.Detection snippet
For anyone else hitting this —
systemctl list-units 'hiddify*'will look fine, these two will not:# units systemd still reports active although the unit file is gone for u in $(systemctl list-units 'hiddify*' --no-legend --plain | awk '{print $1}'); do [ "$(systemctl show -p LoadState --value "$u")" = not-found ] && echo "STALE $u" done # dangling unit symlinks for f in /etc/systemd/system/hiddify-*.service; do [ -e "$(readlink -f "$f")" ] || echo "DANGLING $f -> $(readlink "$f")" done
Also worth knowing for anyone upgrading: the upgrade regenerates every proxy path, so any client still holding a pre-upgrade profile gets a redirect to the decoy site on every outbound and hangs on "Connecting" until the subscription is re-pulled.
One more thing from the same upgrade: the bare-IP certificate silently fell back to self-signed
The ACME run during the upgrade failed for the IP identifier:
_identifiers='{"type":"ip","value":"<server-ip>"}' payload='{"identifiers": [{"type":"ip","value":"<server-ip>"}],"profile": "shortlived"}' ... Signing failed: _on_issue_errThe panel then fell back to its self-signed placeholder for that domain entry (a cert whose issuer equals its subject), and nothing surfaced the failure. The generated configs for the IP entry do not set
allowInsecure, so every IP-based config fails certificate verification client-side from that point on, while the domain-based entries keep working — which makes it look like a client or filtering problem rather than a cert problem.Re-running
services/acme.sh/run.shby hand issued the real short-lived IP cert on the first try, so it looks transient rather than a misconfiguration. Two suggestions:- Treat a fall-back-to-self-signed as a warning in the panel / install output instead of a silent substitution, since the symptom shows up much later and looks unrelated.
- Retry the IP identifier once before falling back, given the short-lived profile has a much tighter renewal cadence than the 90-day DNS certs.
Root cause and fix for §3,
hiddify-core-custom-naive-h2:quic protocol error:haproxy's
in-httpmodefrontend addsAlt-Svc: h3=":443"to every response, including the200that opens naive's CONNECT tunnel. hiddify-core's naive client (4.0.4 and 4.1.0) follows it even though the outbound has"quic": false, and it switches later streams to QUIC. The h3 listener can't carry the tunnel. The first requests succeed over h2, and every stream after that fails within a few ms withstream failed: quic protocol error.The server itself is fine: the official NaiveProxy client and upstream sing-box (1.13.0-rc.4, 1.13.1, 1.14.3) all work with the same outbound JSON.
Workaround until the PR is merged: in
/opt/hiddify-manager/.venv313/lib/python3.13/site-packages/hiddifypanel/proxy_v3/proxy_templates/haproxy/server/fronts/in_httpmode.j2:- http-after-response set-header Alt-Svc 'h3=":443"; ma=2592000' + http-after-response set-header Alt-Svc 'h3=":443"; ma=2592000' unless METH_CONNECT
Then run
systemctl restart hiddify-paneland apply configs (or runbash /opt/hiddify-manager/scripts/apply_configs.sh). Regular web responses still advertise h3.For §1 (
migrate_layout.shdisabling the current units): that's fixed on Hiddify-Managerdev(7ce3eab). The workaround is in #5578.
Version: 13.0.3, upgraded from 12.3.3 (non-docker, Ubuntu 24.04)
1. The old
hiddify-singboxprocess is not stopped, sohiddify-corecrash-loopsAfter the upgrade,
hiddify-singbox.servicewas still running as anot-found activeunit: its unit file was removed, but the process from v12 was never stopped. It still held127.0.0.1:2000, sohiddify-corerestarted more than 100 times with:hiddify-panel-background-tasks,hiddify-dnstm-routerandhiddify-dnstt-<domain>were in the same state: still running, with dangling unit links pointing into the moved v12 folders.Workaround:
systemctl stop hiddify-singbox hiddify-panel-background-tasks, then remove the dangling links in/etc/systemd/system.Suggestion: in
rename_singbox_unit, runsystemctl stop/killon the unit even when the unit file is already missing, and do the same for every unit that has been removed.2.
systemctl reload hiddify-haproxycrashes haproxy 3.4 (awslc)A reload right after certificates were replaced in
data/ssl:systemd restarted it (so there was a short outage). Exit code 132 looks like SIGILL in the new worker. The scripts call
systemctl reload hiddify-haproxyin several places (acme install/run,stop_nginx_acme).3. Two proxy presets don't work after a clean apply (tested with a sing-box-based client)
hiddify-core-rawhttp-vless-http/hiddify-core-rawhttp-vmess-httpon :80:v2ray-http: unexpected status: 404 Not Found. The xray variants of the same presets work.hiddify-core-custom-naive-h2:quic protocol error/ TLS handshakenet_error -100.TUIC, Hysteria, Hysteria2, Mieru, SS, SSH and all the WS/gRPC/HTTPUpgrade/Reality presets work on the same server.