Skip to content

builders: cap parallel build steps and detect frozen builders - #54

Open
ismoilovdevml wants to merge 3 commits into
mainfrom
fix/builder-overload
Open

ismoilovdevml wants to merge 3 commits into
mainfrom
fix/builder-overload

Conversation

@ismoilovdevml

Copy link
Copy Markdown
Owner

Problem

On the trial host, a pipeline that builds ten .NET images at once failed all its build jobs, twice.

  • All ten builds ran in the project's single builder (8 GB, 16 vCPU). buildkitd had no parallelism cap.
  • When the builds reached RUN dotnet build, the guest's memory filled. The guest did not kill a process; it thrashed the page cache instead. Firecracker metrics showed about 13 GB/min of block reads and 0 writes, and sshd did not answer within 90 s.
  • The builder probe only opened a TCP connection. The guest kernel still accepts those while buildkitd never runs, so the frozen builder stayed ready.
  • Busy builders were never dropped. The jobs waited about 23 minutes for their connections to time out.

Change

  • builder.max_parallelism sets [worker.oci] max-parallelism in buildkitd.toml. 0 derives the cap as one step per 6 GB of builder.memory_mb.
    • An explicit value is part of the builder spec. The new spec form replaces every builder once, after it has been idle for 10 minutes, with its cache saved.
    • 0 is not written to config.yaml, so a rollback binary still loads it.
  • The probe completes a TLS handshake against the builder's own CA. An alert from buildkitd counts as an answer, because the probe sends no client certificate. A record without a usable CA keeps the connect check.
  • A busy builder that has not answered for 5 minutes is dropped, regardless of daemon.reconcile_interval. Only the project's next builds benefit. Builds still connected to the frozen builder wait for their connection to time out, and the docs say so.
  • Builders from the previous run are probed in parallel at startup.
  • The docs and the FireRunnerBuildersFailing alert now point at builder.max_parallelism instead of builder.max.

Evidence (KVM, 16 vCPU / 8 GB microVM, buildkitd v0.33.0)

Run Result
10 builds, max-parallelism = 4 The cap held: at most 4 dotnet steps ran at once, across all sessions. The guest still froze at dotnet build (16 GB/min reads, ssh 159 s).
1 build, 16 CPUs Peak 6.3 GB, 277 s
1 build, limited to 4 CPUs Peak 4.8 GB, 270 s

These numbers set the derived cap to one step per 6 GB. 16 vCPUs gave no speedup.

Deployed on the trial host

The trial host runs the same commits on top of the release it was already running.

  • All 5 builders were adopted through the TLS probe.
  • No warnings over three reconciles.
  • No TLS noise in the buildkitd log.
  • The idle builders were replaced as config changed, with their caches saved.

Tests

  • go test ./... -race -count=1, golangci-lint (0 issues) and promtool test rules all pass.
  • The new tests drive a real tls.Listen (TLS 1.2 and 1.3, client certificate required). They also cover a frozen listener (the kernel accepts, nothing is served), a non-TLS service, a closed port, another builder's certificate and no CA.
  • QA mutation run: 13/13 mutants killed. The surviving CA-propagation gap is now covered by TestBuilderProbesUseBuilderCA.

Not solved here

A per-project builder with fixed memory caps how many heavy builds a project can run at once, and it holds its memory while idle. Building in the job's own microVM is the follow-up design.

All 20 build jobs of a 10-image .NET pipeline failed on the trial host:
the ten builds ran `dotnet build` at once in the project's one 8 GB
builder, the guest's memory filled, and it thrashed the page cache
(13 GB/min of disk reads, no writes) instead of killing a process.
The probe only opened a TCP connection, which the guest kernel accepts
while buildkitd never runs, so the frozen builder stayed "ready" and its
jobs waited ~23 minutes for their connections to time out.

- buildkitd.toml sets [worker.oci] max-parallelism from the new
  builder.max_parallelism (0: one step per 2 GB of builder.memory_mb).
  The explicit value is part of the builder spec, so idle builders are
  replaced once (caches saved); a new size alone still replaces none.
  0 is not written to config.yaml, so a rollback binary still loads it.
- The builder probe completes a TLS handshake against the builder's own
  CA (an alert from buildkitd counts as an answer); a record without a
  usable CA keeps the connect check.
- A busy builder silent for 5 reconciles is dropped, so its jobs fail
  then instead of after their TCP timeout.
- Docs and the FireRunnerBuildersFailing alert point at
  builder.max_parallelism instead of builder.max.
Review follow-ups for the builder overload fix:
- A busy builder is dropped after builderBusySilence (5 min) without an
  answer instead of 5 reconciles, so daemon.reconcile_interval 10s no
  longer drops a heavily loaded builder after ~50 s. An answer starts
  the silence over.
- Test that reconcile and adoption hand the builder's own CA to the
  probe (without it the probe falls back to the TCP connect).
- Adoption probes the previous run's builders all at once; each silent
  one can now take two dials and two handshakes.
- Docs say only what holds: the cap makes builds take turns; dropping a
  frozen builder gives the next builds a new one, while builds still
  connected to it wait for their connection to time out; every builder
  is replaced once after the upgrade.
The 2 GB per step guess let an 8 GB builder run four steps at once. On
KVM, four parallel `dotnet build` steps of a 60-project solution froze
the 8 GB / 16 vCPU builder the same way as no cap: a single such step
peaked at 6.3 GB (4.8 GB on 4 vCPUs, same build time). One step per
6 GB keeps an 8 GB builder to one heavy step at a time.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant