Repository navigation
builders: cap parallel build steps and detect frozen builders - #54
Open
ismoilovdevml wants to merge 3 commits into
Open
ismoilovdevml wants to merge 3 commits into
ismoilovdevml wants to merge 3 commits into
Conversation
All 20 build jobs of a 10-image .NET pipeline failed on the trial host: the ten builds ran `dotnet build` at once in the project's one 8 GB builder, the guest's memory filled, and it thrashed the page cache (13 GB/min of disk reads, no writes) instead of killing a process. The probe only opened a TCP connection, which the guest kernel accepts while buildkitd never runs, so the frozen builder stayed "ready" and its jobs waited ~23 minutes for their connections to time out. - buildkitd.toml sets [worker.oci] max-parallelism from the new builder.max_parallelism (0: one step per 2 GB of builder.memory_mb). The explicit value is part of the builder spec, so idle builders are replaced once (caches saved); a new size alone still replaces none. 0 is not written to config.yaml, so a rollback binary still loads it. - The builder probe completes a TLS handshake against the builder's own CA (an alert from buildkitd counts as an answer); a record without a usable CA keeps the connect check. - A busy builder silent for 5 reconciles is dropped, so its jobs fail then instead of after their TCP timeout. - Docs and the FireRunnerBuildersFailing alert point at builder.max_parallelism instead of builder.max.
Review follow-ups for the builder overload fix: - A busy builder is dropped after builderBusySilence (5 min) without an answer instead of 5 reconciles, so daemon.reconcile_interval 10s no longer drops a heavily loaded builder after ~50 s. An answer starts the silence over. - Test that reconcile and adoption hand the builder's own CA to the probe (without it the probe falls back to the TCP connect). - Adoption probes the previous run's builders all at once; each silent one can now take two dials and two handshakes. - Docs say only what holds: the cap makes builds take turns; dropping a frozen builder gives the next builds a new one, while builds still connected to it wait for their connection to time out; every builder is replaced once after the upgrade.
The 2 GB per step guess let an 8 GB builder run four steps at once. On KVM, four parallel `dotnet build` steps of a 60-project solution froze the 8 GB / 16 vCPU builder the same way as no cap: a single such step peaked at 6.3 GB (4.8 GB on 4 vCPUs, same build time). One step per 6 GB keeps an 8 GB builder to one heavy step at a time.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
On the trial host, a pipeline that builds ten .NET images at once failed all its build jobs, twice.
buildkitdhad no parallelism cap.RUN dotnet build, the guest's memory filled. The guest did not kill a process; it thrashed the page cache instead. Firecracker metrics showed about 13 GB/min of block reads and 0 writes, and sshd did not answer within 90 s.buildkitdnever runs, so the frozen builder stayedready.Change
builder.max_parallelismsets[worker.oci] max-parallelisminbuildkitd.toml.0derives the cap as one step per 6 GB ofbuilder.memory_mb.0is not written toconfig.yaml, so a rollback binary still loads it.buildkitdcounts as an answer, because the probe sends no client certificate. A record without a usable CA keeps the connect check.daemon.reconcile_interval. Only the project's next builds benefit. Builds still connected to the frozen builder wait for their connection to time out, and the docs say so.FireRunnerBuildersFailingalert now point atbuilder.max_parallelisminstead ofbuilder.max.Evidence (KVM, 16 vCPU / 8 GB microVM, buildkitd v0.33.0)
max-parallelism = 4dotnetsteps ran at once, across all sessions. The guest still froze atdotnet build(16 GB/min reads, ssh 159 s).These numbers set the derived cap to one step per 6 GB. 16 vCPUs gave no speedup.
Deployed on the trial host
The trial host runs the same commits on top of the release it was already running.
buildkitdlog.config changed, with their caches saved.Tests
go test ./... -race -count=1,golangci-lint(0 issues) andpromtool test rulesall pass.tls.Listen(TLS 1.2 and 1.3, client certificate required). They also cover a frozen listener (the kernel accepts, nothing is served), a non-TLS service, a closed port, another builder's certificate and no CA.TestBuilderProbesUseBuilderCA.Not solved here
A per-project builder with fixed memory caps how many heavy builds a project can run at once, and it holds its memory while idle. Building in the job's own microVM is the follow-up design.