Skip to content

Reaper connection carries no heartbeat, so a NAT that drops idle TCP reaps containers mid-run #3936

Description

@3li7alaki

testcontainers-go: reaper connection carries no heartbeat, so a NAT that drops idle TCP reaps containers mid-run

What happens. Reaper.connect (reaper.go, v0.41.0 to v0.44.0) dials ryuk, sends the
label filter, reads the ACK, and then holds the connection idle for the whole session. On
runtimes whose port forwarding drops idle connections (ArcBox measured at 150s, arcboxlabs/arcbox#732),
ryuk sees the client disconnect and reaps every container RYUK_RECONNECTION_TIMEOUT (10s)
later, while tests are still using them. It shows up as postgres killed with exit 137 a few
minutes into a suite.

Workaround today. Raise RYUK_RECONNECTION_TIMEOUT (we set 30m on VM-backed runtimes
only). That trades the mid-run kill for a long cleanup delay after a crash, and it still
fails for sessions longer than the timeout.

Proposal. Re-send the filter line on the open connection on an interval below common NAT
idle timeouts (say every 30s). Ryuk already accepts repeated filter lines and ACKs each one,
so this needs no ryuk change and works on every runtime, and the default 10s reconnection
timeout stays safe.

Activity

  1. Yusuf-Hussien commented on Oct 7, 2026

    @Yusuf-Hussien

    Confirmed on main. The analysis is accurate, and the failure mechanism is easy to trace in reaper.go:

    • Reaper.Connect (reaper.go:470) ظ�ْ connect (reaper.go:538) dials ryuk and spawns the reader goroutine (reaper.go:546-552).
    • handshake (reaper.go:557) writes the one-time label filter and reads the single ACK, then returns.
    • After handshake returns, the goroutine goes straight to <-terminationSignal (reaper.go:551) and the TCP connection sits idle for the entire test session. No keep-alive, no periodic traffic.

    The proposal is solid and needs no ryuk change: ryuk treats each received line as a filter update and ACKs it, so re-sending the same filter on an interval proves liveness to the underlying runtime's port-forwarding layer. Two small implementation notes for the fix:

    1. The heartbeat interval must be configurable / coordinated with the existing RYUK_RECONNECTION_TIMEOUT and RYUK_CONNECTION_TIMEOUT handling at reaper.go:408-411 (heartbeat period must be well below the NAT idle timeout, e.g. every 30s as suggested).
    2. After each heartbeat write we should consume ryuk's ACK response (or the read side could fill over a long session), reusing the existing ACK-read logic from handshake rather than duplicating the protocol.

    This makes the default 10s reconnection timeout safe again and removes the RYUK_RECONNECTION_TIMEOUT bump workaround for long suites without a crash-cleanup delay trade-off.

    I'd like to take this: add a heartbeat ticker in the connect goroutine that re-sends the filter line and drains the ACK, with a test in reaper_test.go that asserts the connection stays alive across an interval longer than the current idle timeout.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions