Skip to content

v13.0.3 upgrade from v12: old hiddify-singbox keeps port 2000 so hiddify-core crash-loops; haproxy 3.4 crashes on reload; rawhttp-HC/naive-h2 presets fail #5566

Description

@arnevis

Version: 13.0.3, upgraded from 12.3.3 (non-docker, Ubuntu 24.04)

1. The old hiddify-singbox process is not stopped, so hiddify-core crash-loops

After the upgrade, hiddify-singbox.service was still running as a not-found active unit: its unit file was removed, but the process from v12 was never stopped. It still held 127.0.0.1:2000, so hiddify-core restarted more than 100 times with:

FATAL[0000] start service: manager start inbound/socks[socks-in]: listen tcp 127.0.0.1:2000: bind: address already in use

hiddify-panel-background-tasks, hiddify-dnstm-router and hiddify-dnstt-<domain> were in the same state: still running, with dangling unit links pointing into the moved v12 folders.
Workaround: systemctl stop hiddify-singbox hiddify-panel-background-tasks, then remove the dangling links in /etc/systemd/system.
Suggestion: in rename_singbox_unit, run systemctl stop / kill on the unit even when the unit file is already missing, and do the same for every unit that has been removed.

2. systemctl reload hiddify-haproxy crashes haproxy 3.4 (awslc)

A reload right after certificates were replaced in data/ssl:

[ALERT] exit-on-failure: killing every processes with SIGTERM
[WARNING] All workers exited. Exiting... (132)
hiddify-haproxy.service: Main process exited, code=exited, status=132/n/a

systemd restarted it (so there was a short outage). Exit code 132 looks like SIGILL in the new worker. The scripts call systemctl reload hiddify-haproxy in several places (acme install/run, stop_nginx_acme).

3. Two proxy presets don't work after a clean apply (tested with a sing-box-based client)

  • hiddify-core-rawhttp-vless-http / hiddify-core-rawhttp-vmess-http on :80: v2ray-http: unexpected status: 404 Not Found. The xray variants of the same presets work.
  • hiddify-core-custom-naive-h2: quic protocol error / TLS handshake net_error -100.

TUIC, Hysteria, Hysteria2, Mieru, SS, SSH and all the WS/gRPC/HTTPUpgrade/Reality presets work on the same server.

Activity

  1. rezasabourinejad commented on Sep 27, 2026

    @rezasabourinejad

    Same upgrade path here (13.0.3 from 12.x, non-docker, Ubuntu 24.04). I can confirm §1, and I think I found the exact root cause — plus a second, independent cause of the same address already in use symptom.

    Root cause of the stale units: the cleanup guards use [ -e ] on a symlink

    All three teardown blocks for legacy units are guarded the same way:

    • scripts/migrate_layout.sh:50 (rename_singbox_unit) — [ -e /etc/systemd/system/hiddify-singbox.service ]
    • services/hiddify-core/run.sh:6 — same test, same unit
    • services/panel/install.sh:29 — [ -e /etc/systemd/system/hiddify-panel-background-tasks.service ]

    /etc/systemd/system/hiddify-*.service is a symlink into the versioned tree (e.g. -> /opt/hiddify-manager/hiddify-panel/hiddify-panel-background-tasks.service). [ -e ] dereferences the link, so once the same upgrade has already deleted the old directory the test is false, and the whole block — the systemctl disable --now and the rm -f — is skipped. The guard fails in exactly the situation it was written for.

    One-line fix in all three places:

    if [ -e /etc/systemd/system/<unit> ] || [ -L /etc/systemd/system/<unit> ]; then

    That is also why this is easy to miss: since the rm -f never runs, the dangling link stays and systemd keeps the unit as LoadState=not-found with ActiveState=active, so systemctl status and systemctl list-units both look healthy while a v12 process is still running.

    The hiddify-panel-background-tasks case specifically

    Besides the singbox process holding 127.0.0.1:2000, the v12 celery worker+beat pair from the deleted hiddify-panel/ was still running after 29 days, executing pre-MySQL code against the new MySQL-backed panel. Symptoms:

    • MySQLdb OperationalError (2013, 'Lost connection to MySQL server during query') in hiddify_panel.err.log every few minutes
    • every user's last_online and current_usage frozen at the upgrade timestamp — the panel reported 0 users online while ~780 established proxy connections and ~0.5 MiB/s were actually flowing

    Killing the two PIDs and removing the dangling link fixed it. The in-process APScheduler then picked accounting back up on its own at roughly a one-minute interval, so no cron was needed — the folding-in works, it just never gets a chance while the old worker is still alive.

    Second, independent cause of the same address already in use

    Worth splitting into its own issue if you prefer — it is not upgrade-specific and should be reproducible on a clean install.

    After clearing the orphan, hiddify-core still crash-looped, failing on a different port every restart (49578, 49886, 49138, 48412, …). The generated inbound ports for both xray.json and hiddify-core.json sit inside the kernel's default ephemeral range:

    net.ipv4.ip_local_port_range     = 32768  60999
    net.ipv4.ip_local_reserved_ports =            # empty
    

    27 of my 31 hiddify-core inbound ports are inside that range. So any outbound socket on the box can be handed one of those ports as its source port, and hiddify-core then loses the race on startup. scripts/common/sysctl.conf tunes a lot of TCP knobs but never touches either of these two.

    The actual squatter turned out to be rpxy-l4: it leaks CLOSE-WAIT sockets to its own target (127.0.0.1:<ephemeral> -> 127.0.0.1:901) — about 970 of them at the time. CLOSE-WAIT has no kernel timeout, so those ports stay taken until the process is restarted.

    Two separate things would help:

    1. Add the generated inbound ports to net.ipv4.ip_local_reserved_ports in scripts/common/sysctl.conf, or allocate them below 32768. Note that reserving only blocks future automatic allocation, so on an upgrade a process that already leaked into the range still has to be restarted once.
    2. The rpxy-l4 CLOSE-WAIT leak looks like a socket-lifetime bug in its own right.

    What I am running as a workaround:

    # /etc/sysctl.d/99-hiddify-reserved-ports.conf
    net.ipv4.ip_local_reserved_ports = <generated inbound ports, e.g. 47900-49960>
    

    then restarting rpxy-l4 once to drop the already-leaked sockets before starting hiddify-core.

    Detection snippet

    For anyone else hitting this — systemctl list-units 'hiddify*' will look fine, these two will not:

    # units systemd still reports active although the unit file is gone
    for u in $(systemctl list-units 'hiddify*' --no-legend --plain | awk '{print $1}'); do
      [ "$(systemctl show -p LoadState --value "$u")" = not-found ] && echo "STALE $u"
    done
    
    # dangling unit symlinks
    for f in /etc/systemd/system/hiddify-*.service; do
      [ -e "$(readlink -f "$f")" ] || echo "DANGLING $f -> $(readlink "$f")"
    done

    Also worth knowing for anyone upgrading: the upgrade regenerates every proxy path, so any client still holding a pre-upgrade profile gets a redirect to the decoy site on every outbound and hangs on "Connecting" until the subscription is re-pulled.

    One more thing from the same upgrade: the bare-IP certificate silently fell back to self-signed

    The ACME run during the upgrade failed for the IP identifier:

    _identifiers='{"type":"ip","value":"<server-ip>"}'
    payload='{"identifiers": [{"type":"ip","value":"<server-ip>"}],"profile": "shortlived"}'
    ...
    Signing failed:
    _on_issue_err
    

    The panel then fell back to its self-signed placeholder for that domain entry (a cert whose issuer equals its subject), and nothing surfaced the failure. The generated configs for the IP entry do not set allowInsecure, so every IP-based config fails certificate verification client-side from that point on, while the domain-based entries keep working — which makes it look like a client or filtering problem rather than a cert problem.

    Re-running services/acme.sh/run.sh by hand issued the real short-lived IP cert on the first try, so it looks transient rather than a misconfiguration. Two suggestions:

    1. Treat a fall-back-to-self-signed as a warning in the panel / install output instead of a silent substitution, since the symptom shows up much later and looks unrelated.
    2. Retry the IP identifier once before falling back, given the short-lived profile has a much tighter renewal cadence than the 90-day DNS certs.
  2. blshkv commented on Oct 9, 2026

    @blshkv

    Root cause and fix for §3, hiddify-core-custom-naive-h2: quic protocol error:

    haproxy's in-httpmode frontend adds Alt-Svc: h3=":443" to every response, including the 200 that opens naive's CONNECT tunnel. hiddify-core's naive client (4.0.4 and 4.1.0) follows it even though the outbound has "quic": false, and it switches later streams to QUIC. The h3 listener can't carry the tunnel. The first requests succeed over h2, and every stream after that fails within a few ms with stream failed: quic protocol error.

    The server itself is fine: the official NaiveProxy client and upstream sing-box (1.13.0-rc.4, 1.13.1, 1.14.3) all work with the same outbound JSON.

    Workaround until the PR is merged: in /opt/hiddify-manager/.venv313/lib/python3.13/site-packages/hiddifypanel/proxy_v3/proxy_templates/haproxy/server/fronts/in_httpmode.j2:

    -  http-after-response set-header Alt-Svc 'h3=":443"; ma=2592000'
    +  http-after-response set-header Alt-Svc 'h3=":443"; ma=2592000' unless METH_CONNECT

    Then run systemctl restart hiddify-panel and apply configs (or run bash /opt/hiddify-manager/scripts/apply_configs.sh). Regular web responses still advertise h3.

    PR: hiddify/hiddifypanel#46

    For §1 (migrate_layout.sh disabling the current units): that's fixed on Hiddify-Manager dev (7ce3eab). The workaround is in #5578.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions