Skip to content

migrate_layout.sh` disables the *current* core services on every install/apply after a v12 → v13 upgrade (panel, redis, nginx, haproxy go down; "Unit ... not found") #5562

Description

@Survivor-viper

Bug/Feature:

Description: Describe bug or needed feature

Details:

Hiddify Version: 13.0.3
Python Version: 3.13.15 (main, Aug 25 2026, 14:01:07) [Clang 22.1.3 ]
OS: Linux-6.8.0-31-generic-x86_64-with-glibc2.39
User Agent: Unknown

migrate_layout.sh disables the current core services on every install/apply after a v12 → v13 upgrade (panel, redis, nginx, haproxy go down; "Unit ... not found")

Version: HiddifyManager 13.0.3 (upgraded automatically from 12.x), Ubuntu, non-docker install.

Symptoms

After the automatic upgrade to 13.0.3, and then again after every "Apply configs" (admin panel button, scripts/apply_configs.sh, scheduled apply):

  • admin panel is unreachable, all TLS/QUIC profiles on 443 stop working (only REALITY served by xray keeps working);
  • hiddify-panel, hiddify-redis, hiddify-nginx, hiddify-haproxy (and mtproxy) are gone from systemd;
  • the install/apply log contains:
Failed to start hiddify-redis.service: Unit hiddify-redis.service not found.
redis.exceptions.ConnectionError: Error 111 connecting to 127.0.0.1:6379. Connection refused.
Failed to start hiddify-panel.service: Unit hiddify-panel.service not found.
Failed to dump server configs into /opt/hiddify-manager/generated
Failed to restart hiddify-nginx.service: Unit hiddify-nginx.service not found.
Failed to start hiddify-haproxy.service: Unit hiddify-haproxy.service not found.

A full scripts/install.sh install-production brings everything back (because each services/*/install.sh re-creates the unit symlink), but the next apply kills it again. That is why the problem looks random and keeps coming back.

Root cause

scripts/migrate_layout.sh runs at the start of every scripts/install.sh call (install.sh lines 5–6), including apply_configs (DO_NOT_INSTALL=true).

Once the migration is done, singbox/ no longer exists, so the "already migrated" branch runs every time (migrate_layout.sh lines 84–96) and it calls disable_all_services:

# migrate_layout.sh, line 43
disable_all_services() {
    find /opt/hiddify-manager/ -type d -name "services" -prune -o -type f -name "*.service" -print |
    xargs -r -n1 basename |
    tee /dev/stderr |
    xargs -r -I{} systemctl disable --now {}
}

The find only prunes directories named services. It does descend into /opt/hiddify-manager/old/, where the migration itself moved the v12 unit files. Those files have the same names as the new units:

old/hiddify-panel/hiddify-panel.service
old/other/redis/hiddify-redis.service
old/nginx/hiddify-nginx.service
old/haproxy/hiddify-haproxy.service
old/xray/hiddify-xray.service
old/other/telegram/*/mtproxy.service
...

xargs basename keeps only the name, so systemctl disable --now hiddify-panel.service hits the current unit. The current units are installed as symlinks (ln -sf services/<x>/hiddify-<x>.service /etc/systemd/system/). For such linked units, systemctl disable also removes the symlink, so the unit disappears ("not found").

In a full install, install_run → services/<x>/install.sh re-creates the link. In an apply (DO_NOT_INSTALL=true), install.sh is skipped (install_run, if [ "$DO_NOT_INSTALL" != "true" ]), so the services stay deleted and the following run.sh calls fail.

Steps to reproduce

  1. Have a 12.x installation, let it auto-upgrade to 13.0.3, or upgrade manually. old/ gets populated.
  2. Run bash /opt/hiddify-manager/scripts/apply_configs.sh --no-gui, or press Apply in the panel.
  3. systemctl status hiddify-panel → Unit hiddify-panel.service could not be found.

Check what the function will hit:

find /opt/hiddify-manager/ -type d -name "services" -prune -o -type f -name "*.service" -print

Any output from under old/ with a name that also exists under services/ means the current unit will be disabled.

Proposed fix

1. scripts/migrate_layout.sh: never touch old/, and never disable a unit name that the new layout ships.

 disable_all_services() {
-    find /opt/hiddify-manager/ -type d -name "services" -prune -o -type f -name "*.service" -print |
-    xargs -r -n1 basename |
-    tee /dev/stderr |
-    xargs -r -I{} systemctl disable --now {}
+    # old/ holds v12 copies with the SAME unit names as services/
+    # (hiddify-panel.service, hiddify-redis.service, ...). Disabling them by
+    # name would disable, and for linked units DELETE, the current units.
+    local current
+    current="$(find "$SERVICES" -type f -name '*.service' -printf '%f\n' | sort -u)"
+    find "$HIDDIFY_DIR" \( -path "$SERVICES" -o -path "$OLD" \) -prune \
+        -o -type f -name '*.service' -printf '%f\n' |
+    sort -u |
+    grep -vxF -e "$current" |
+    tee /dev/stderr |
+    xargs -r -I{} systemctl disable --now {}
 }

($SERVICES, $OLD and $HIDDIFY_DIR are already defined at the top of the script. grep -e with a newline-separated string treats every line as a separate fixed pattern.)

Also consider calling disable_all_services only when a migration actually happens (the singbox/ found branch), not on every run in the "already migrated" branch at line 85. Disabling services is not something an apply should do.

2. Defensive: scripts/install.sh: an apply should re-create missing unit links instead of assuming they exist.

function ensure_unit_links() {
    local changed=0 u name
    for u in /opt/hiddify-manager/$1/*.service; do
        [ -f "$u" ] || continue
        name="$(basename "$u")"
        case "$name" in *@.service) continue ;; esac   # template units are linked by their own install.sh
        if [ ! -e "/etc/systemd/system/$name" ]; then
            ln -sf "$u" "/etc/systemd/system/$name"
            changed=1
        fi
    done
    if [ "$changed" = 1 ]; then systemctl daemon-reload; fi
}

function install_run() {
    echo "======================$1====================================={"
    local start_time=$(date +%s)

    if [ "$DO_NOT_INSTALL" != "true" ];then
            runsh install.sh $@
        if [ "$MODE" != "apply_users" ] && [ "$DOCKER_MODE" != "true"  ]; then
            systemctl daemon-reload
        fi
    elif [ "$DOCKER_MODE" != "true" ]; then
        ensure_unit_links "$1"
    fi
    ...

With fix 1 alone the bug is gone. Fix 2 makes apply survive any other future cause of a missing link.

Workaround for affected servers

# 1. hide the old unit files from disable_all_services (files are kept, only renamed)
find /opt/hiddify-manager/old -type f -name '*.service' -exec mv {} {}.disabled \;
# 2. re-create the links once
cd /opt/hiddify-manager && bash scripts/install.sh install-production --no-gui

We verified the workaround on 13.0.3: before it, every apply removed hiddify-panel/redis/nginx/haproxy. After renaming the old unit files, an apply finishes with all services active and the panel reachable.


Related (separate issue): firewall template opens legacy ports, not the ports v13 proxies listen on

services/firewall/run.sh.j2 (lines 27–38) opens UDP ports from the legacy per-domain fields internal_port_hysteria2, internal_port_tuic and internal_port_naive:

{% for d in domains if d['internal_port_hysteria2'] or d['internal_port_tuic'] or d['internal_port_naive']  %}
  {% if d['internal_port_hysteria2']>0 %}
  allow_port "udp" {{d['internal_port_hysteria2']}} #hysteria2
  ...

In 13.x the inbound ports come from custom_proxy.server_inbound_tcp_ports and custom_proxy.server_inbound_udp_ports. After the migration, Hysteria2 (hiddify-core-custom-hysteria2-quic) got UDP 33512, while the firewall still opened the old 39440/39443. The generated hiddify-core.json listened on 33512, but the default UFW reject chain dropped it, so Hysteria2 silently never worked.

Suggested fix: build the allow list from the enabled mode == ip custom proxies, the same source the core config uses (ProxyRenderCache.proxies[*].server_inbound_tcp_ports / server_inbound_udp_ports with enable == True), instead of domains[*].internal_port_*. For example, pass custom_proxy_ports into the firewall template context:

{% for p in custom_proxy_ports if p.enable %}
  {% for port in p.server_inbound_udp_ports %}allow_port "udp" {{port}} # {{p.server_tag}}
  {% endfor %}
  {% for port in p.server_inbound_tcp_ports %}allow_port "tcp" {{port}} # {{p.server_tag}}
  {% endfor %}
{% endfor %}

(The variable name is illustrative. It should come from the same ProxyRenderCache the core config is rendered from.)


Related (separate issue): hiddify-core.json embeds certificates inline and is not re-rendered after ACME renewal, so Hysteria2 serves a stale or self-signed certificate

proxy_templates/hiddify-core/server/tls/tls.j2 puts the certificate content into the config:

  {% if cert.cert_lines %}
  "certificate": {{ cert.cert_lines | tojson }},
  "key": {{ cert.key_lines | tojson }},
  {% endif %}

The config is rendered by scripts/common/replace_variables.sh → dump_server_configs. That step runs before install_run services/acme.sh in scripts/install.sh, and the ACME paths never render it again:

  • services/acme.sh/run.sh ends with sync_tls_store + systemctl reload hiddify-core. The reload re-reads the same JSON with the old embedded certificate.
  • sync-tls-store (CLI and admin/sync-tls-store/) only imports files into tls_store. It does not dump configs.
  • scripts/common/daily_actions.sh (cron @daily) only runs services/acme.sh/run.sh.

Timeline from one normal apply on 13.0.3 (mtimes):

07:24:50  generated/hiddify-core.json rendered (certificate embedded from tls_store)
07:25:11  acme.sh --installcert rewrites data/ssl/<server-ip>.crt, sync_tls_store, reload hiddify-core

So after every apply the core serves the certificate that existed before the ACME step. IP certificates are issued with --certificate-profile shortlived --days 6 (cert_utils.sh, get_cert). The daily renewal therefore updates the files and tls_store, but Hysteria2 keeps serving the embedded one. Once the embedded one expires, every Hysteria2 client fails until someone presses Apply. haproxy is not affected because it reads data/ssl/ files.

In our case it was worse. A broken apply (see the first issue) had left the temporary placeholder from get_self_signed_cert (O=Google Trust Services LLC, self-signed) embedded in hiddify-core.json, while haproxy on 443 already served the Let's Encrypt one. Hysteria2 clients failed with:

CRYPTO_ERROR 0x12a (local): tls: failed to verify certificate

Re-running hiddify-panel-cli dump-server-configs + systemctl reload hiddify-core (no other change) fixed it immediately.

Proposed fix, either one:

A. Reference the files, so a reload picks up renewed certificates (sing-box/hiddify-core support certificate_path and key_path; the unit has ExecReload=/bin/kill -HUP $MAINPID):

  {% if cert.cert_path and cert.key_path %}
  "certificate_path": "{{ cert.cert_path }}",
  "key_path": "{{ cert.key_path }}",
  {% elif cert.cert_lines %}
  "certificate": {{ cert.cert_lines | tojson }},
  "key": {{ cert.key_lines | tojson }},
  {% endif %}

CertVar would need cert_path and key_path filled from tls_store_sync.cert_paths_for_domain(domain). The files are group hiddify-common readable; please check the core's user is in that group.

B. Re-render after ACME. services/acme.sh/run.sh already sources scripts/common/utils.sh, where dump_server_configs is defined:

 set_files_in_folder_readable_to_hiddify_common_group /opt/hiddify-manager/data/ssl
 # Refresh tls_store after batch ACME / self-signed so DB matches installed files.
 sync_tls_store || true
+# hiddify-core.json embeds certificates inline: re-render, or reload keeps the old ones
+dump_server_configs || true
 systemctl reload hiddify-haproxy
 systemctl reload hiddify-core

Related (separate issue): a failed apply deletes all TLS certificates, and a glob bug then creates cert_utils.sh.crt

scripts/common/replace_variables.sh deletes every data/ssl/*.crt whose name is not in current.json:

domains=$(cat /opt/hiddify-manager/data/current.json | jq -r '.domains[] | .domain' | tr '\n' ' ')
for f in /opt/hiddify-manager/data/ssl/*.crt; do
    d=$(basename "$f" .crt)
    if [[ ! " ${domains[@]} " =~ " ${d} " ]]; then
        rm "/opt/hiddify-manager/data/ssl/$d.crt"
        rm "/opt/hiddify-manager/data/ssl/$d.crt.key"
    fi
done

When the panel is down (redis refused, as in the first issue), the domain list comes out empty, so all certificates, including valid Let's Encrypt ones, are deleted. The next steps replace them with self-signed placeholders.

services/acme.sh/run.sh then iterates over an empty directory:

for f in /opt/hiddify-manager/data/ssl/*.crt; do
    d=$(basename "$f" .crt)      # no match -> d="*"
    get_self_signed_cert $d &    # unquoted -> glob in cwd (services/acme.sh/) -> "cert_utils.sh"
done

The log of that apply shows exactly this:

Certificate cert_utils.sh (/opt/hiddify-manager/data/ssl/cert_utils.sh.crt) file not found. Generating a new certificate.

Proposed fix:

-domains=$(cat /opt/hiddify-manager/data/current.json | jq -r '.domains[] | .domain' | tr '\n' ' ')
+domains=$(jq -r '.domains[] | .domain' /opt/hiddify-manager/data/current.json 2>/dev/null | tr '\n' ' ')
+if [ -z "${domains// /}" ]; then
+    echo "current.json has no domains (panel down?) - not deleting any certificate" >&2
+else
 for f in /opt/hiddify-manager/data/ssl/*.crt; do
     ...
 done
+fi

and in services/acme.sh/run.sh: add shopt -s nullglob before the for f in .../*.crt loop, and quote the argument (get_self_signed_cert "$d" &, also for the $domains and $fake_domains loops).


Related (separate issue): migration 12 → 13 silently drops Hysteria2 (and TUIC)

On 12.x Hysteria2 worked with hysteria_enable=1. After the auto-upgrade:

  • the Hysteria2 custom proxy (hiddify-core-custom-hysteria2-quic) has enable=1, but it is effectively disabled by the parent switch quic_enable=0 (CustomProxy.blocked_parent_enables()). TUIC is blocked by tuic_enable + quic_enable. Nothing in the panel UI or the install log says so;
  • its port was changed from the old hysteria_port-derived 39440 to a random 33512, which neither existing clients nor the firewall template (see the firewall issue) know about.

Result: users who relied on Hysteria2 lose it after the upgrade, with no error anywhere.

Suggestion: when migrating, set quic_enable=1 if hysteria_enable or tuic_enable was on in 12.x, and keep the previous ports for migrated proxies. Alternatively, show "blocked by quic_enable" next to the proxy in the admin UI.


Workarounds we applied on 13.0.3 (all verified on the server)

  1. migrate_layout.sh issue: renamed old/**/*.service → *.service.disabled. After that, Apply keeps all services running (checked twice).
  2. Hysteria2: quic_enable=1, set the Hysteria2 custom proxy UDP port back to 39440 (already open in the firewall), and disabled the xhttp-over-QUIC variants we do not use.
  3. Inline certificates: an hourly cron re-renders configs only when some data/ssl/*.crt is newer than generated/hiddify-core.json, then reloads hiddify-core and haproxy:
pgrep -f "scripts/install.sh" >/dev/null && exit 0
find /opt/hiddify-manager/data/ssl -maxdepth 1 -name '*.crt' -newer /opt/hiddify-manager/generated/hiddify-core.json | grep -q . || exit 0
cd /opt/hiddify-manager/services/panel
sudo -u hiddify-panel env HIDDIFY_CFG_PATH=/opt/hiddify-manager/data/hiddify-panel/app.cfg \
    /opt/hiddify-manager/.venv313/bin/python -m hiddifypanel dump-server-configs /opt/hiddify-manager/generated \
  && systemctl reload hiddify-core && systemctl reload hiddify-haproxy

After 2 and 3, Hysteria2 serves the Let's Encrypt certificate and passes traffic, and it still does after another Apply.

Activity

  1. arnevis commented on Sep 27, 2026

    @arnevis

    I hit the same thing (12.3.3 → 13.0.3). One more consequence: with hiddify-redis disabled, hiddify-panel-cli all-configs crashes, so data/current.json ends up empty and replace_variables.sh then deletes every cert in data/ssl. Details in #5565. My workaround for this issue: move /opt/hiddify-manager/old out of the tree, then re-link and enable the units from services/*/.

  2. rezasabourinejad commented on Sep 27, 2026

    @rezasabourinejad

    Confirming this on another server: 13.0.3, auto-upgraded from 12.x by the 03:00 cron, Ubuntu, non-docker.

    In the panel it shows up as: resetting a user (or saving settings) runs install.sh apply_users. That call disables hiddify-redis, hiddify-panel, hiddify-nginx, hiddify-haproxy, hiddify-xray and hiddify-dnstm-router. Clients then get TLS errors on 443, because rpxy-l4 forwards to 127.0.0.1:901 and nothing is listening there. As #5565 describes, data/ssl was also empty afterwards.

    The workaround from the earlier comment works. I also patched the find so the problem can't come back if old/ reappears.

    In scripts/migrate_layout.sh, disable_all_services(), the first line becomes:

    find /opt/hiddify-manager/ -type d \( -name "services" -o -name "old" -o -name ".venv*" \) -prune -o -type f -name "*.service" -print |

    Then:

    mv /opt/hiddify-manager/old /root/hiddify-old-backup
    bash /opt/hiddify-manager/scripts/install.sh --no-gui   # re-create the units

    After that, install.sh apply_users --no-gui leaves every service active.

    Note: 13.0.3 already includes the commit "fix not disabling old services", and the bug is still there. A proper fix would be one of these:

    • run disable_all_services only while the migration is actually happening, not on every call;
    • or disable a unit only when it is not the same name as a unit under services/.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions