Reaper connection carries no heartbeat, so a NAT that drops idle TCP reaps containers mid-run #3936
Description
Activity
Confirmed on
main. The analysis is accurate, and the failure mechanism is easy to trace inreaper.go:Reaper.Connect(reaper.go:470) ظ�ْconnect(reaper.go:538) dials ryuk and spawns the reader goroutine (reaper.go:546-552).handshake(reaper.go:557) writes the one-time label filter and reads the single ACK, then returns.- After
handshakereturns, the goroutine goes straight to<-terminationSignal(reaper.go:551) and the TCP connection sits idle for the entire test session. No keep-alive, no periodic traffic.
The proposal is solid and needs no ryuk change: ryuk treats each received line as a filter update and ACKs it, so re-sending the same filter on an interval proves liveness to the underlying runtime's port-forwarding layer. Two small implementation notes for the fix:
- The heartbeat interval must be configurable / coordinated with the existing
RYUK_RECONNECTION_TIMEOUTandRYUK_CONNECTION_TIMEOUThandling atreaper.go:408-411(heartbeat period must be well below the NAT idle timeout, e.g. every 30s as suggested). - After each heartbeat write we should consume ryuk's ACK response (or the read side could fill over a long session), reusing the existing ACK-read logic from
handshakerather than duplicating the protocol.
This makes the default 10s reconnection timeout safe again and removes the
RYUK_RECONNECTION_TIMEOUTbump workaround for long suites without a crash-cleanup delay trade-off.I'd like to take this: add a heartbeat ticker in the
connectgoroutine that re-sends the filter line and drains the ACK, with a test inreaper_test.gothat asserts the connection stays alive across an interval longer than the current idle timeout.
testcontainers-go: reaper connection carries no heartbeat, so a NAT that drops idle TCP reaps containers mid-run
What happens.
Reaper.connect(reaper.go, v0.41.0 to v0.44.0) dials ryuk, sends thelabel filter, reads the ACK, and then holds the connection idle for the whole session. On
runtimes whose port forwarding drops idle connections (ArcBox measured at 150s, arcboxlabs/arcbox#732),
ryuk sees the client disconnect and reaps every container
RYUK_RECONNECTION_TIMEOUT(10s)later, while tests are still using them. It shows up as postgres killed with exit 137 a few
minutes into a suite.
Workaround today. Raise
RYUK_RECONNECTION_TIMEOUT(we set 30m on VM-backed runtimesonly). That trades the mid-run kill for a long cleanup delay after a crash, and it still
fails for sessions longer than the timeout.
Proposal. Re-send the filter line on the open connection on an interval below common NAT
idle timeouts (say every 30s). Ryuk already accepts repeated filter lines and ACKs each one,
so this needs no ryuk change and works on every runtime, and the default 10s reconnection
timeout stays safe.