Testcontainers version
v0.44.0
Using the latest Testcontainers version?
Yes
Host OS
macOS 26.6.2
Host arch
arm64
Go version
go1.26.5
Docker version
Docker Desktop 29.6.1, API 1.55, 8 GB VM
What happened
reaperSpawner.newReaper creates reaper_<sessionID> using the caller's context. When the daemon is under load, that ContainerCreate request can outlast the caller's deadline. The client gets a context error and never learns the container ID, but the daemon still completes the create, sometimes seconds later.
I confirmed this outside testcontainers with plain POST /containers/create requests, cancelled on the client side after 5–40ms. Every request timed out, yet 7 of 8 still produced a container in state created.
The orphan is never started. For the rest of the session:
- its fixed name blocks creating a new reaper;
lookupContainer (All: true) keeps returning it;
fromContainer waits for a Started log line that never comes.
Every later package binary therefore fails to connect a reaper, and nothing reaps what those binaries leave behind. On one developer machine this built up to three reapers stuck in Created and 172 leaked containers across three days. Starting one of the stuck reapers by hand (docker start) cleaned up its session within about 70s.
This is close to #3867, but it is a different state. #3868 treats created as a reaper that is on its way up and keeps retrying. For a reaper orphaned by an abandoned create, that retrying still ends in the caller's deadline.
Suggestions
- Recognise a
created reaper that has stayed un-started past a short grace period as stale, then remove and recreate it. The grace period would come from the container's Created timestamp.
- Alternatively, or as well, issue the reaper's create and start with
context.WithoutCancel(ctx) plus a timeout the library owns. That way a caller's deadline cannot orphan the reaper halfway through.
Relevant log output
create container: reaper: from container "<id>": wait for reaper <id>: context deadline exceeded
Additional information
The same abandoned-but-completed behaviour applies to ContainerStart. A start that times out on the client side can still bring the container up afterwards.
Related: testcontainers/moby-ryuk#245, where Ryuk exits without pruning when its container listing times out under the same load.
Testcontainers version
v0.44.0
Using the latest Testcontainers version?
Yes
Host OS
macOS 26.6.2
Host arch
arm64
Go version
go1.26.5
Docker version
Docker Desktop 29.6.1, API 1.55, 8 GB VM
What happened
reaperSpawner.newReapercreatesreaper_<sessionID>using the caller's context. When the daemon is under load, thatContainerCreaterequest can outlast the caller's deadline. The client gets a context error and never learns the container ID, but the daemon still completes the create, sometimes seconds later.I confirmed this outside testcontainers with plain
POST /containers/createrequests, cancelled on the client side after 5–40ms. Every request timed out, yet 7 of 8 still produced a container in statecreated.The orphan is never started. For the rest of the session:
lookupContainer(All: true) keeps returning it;fromContainerwaits for aStartedlog line that never comes.Every later package binary therefore fails to connect a reaper, and nothing reaps what those binaries leave behind. On one developer machine this built up to three reapers stuck in
Createdand 172 leaked containers across three days. Starting one of the stuck reapers by hand (docker start) cleaned up its session within about 70s.This is close to #3867, but it is a different state. #3868 treats
createdas a reaper that is on its way up and keeps retrying. For a reaper orphaned by an abandoned create, that retrying still ends in the caller's deadline.Suggestions
createdreaper that has stayed un-started past a short grace period as stale, then remove and recreate it. The grace period would come from the container'sCreatedtimestamp.context.WithoutCancel(ctx)plus a timeout the library owns. That way a caller's deadline cannot orphan the reaper halfway through.Relevant log output
Additional information
The same abandoned-but-completed behaviour applies to
ContainerStart. A start that times out on the client side can still bring the container up afterwards.Related: testcontainers/moby-ryuk#245, where Ryuk exits without pruning when its container listing times out under the same load.