Skip to content

docs deploy: production builds are OOM-failing now — the peak is the Turbopack compile phase, not the prerender; turbopackSourceMaps is a free knob #12683's option set missed #12711

Description

@os-zhuang

Production docs deploys are failing right now — five consecutive production builds OOM-killed. This is the live-outage counterpart to #12683's standing-risk card, and it corrects the attribution in both #12683 and #12698.

Measured (Vercel API + build logs, 2026-08-27)

Vercel's own API on the failed deployment:

errorCode    = "out_of_memory"
buildMachine = { purchaseType: "standard", cores: 4, memory: 8192 }

The peak is the Turbopack COMPILE phase, not the prerender

Both failures die between Creating an optimized production build ... and Compiled successfully — with zero output in between:

11:31:03.206  Creating an optimized production build ...
11:42:26.615  ERROR ... exited (137)          <- 11m23s of silence
11:43:01.851  Creating an optimized production build ...
11:45:03.086  ERROR ... exited (137)          <- 2m of silence

The build never reaches Collecting page data / Generating static pages. So #12683's attribution — "the docs site prerenders 400+ MDX pages", offered as the reason one process holds 7.6 GB — cannot be the mechanism: that phase does not execute. Same for the 403 build-time OG PNGs (app/og/docs/[...slug]/route.tsx) and the 403 markdown routes (app/llms.mdx/docs/[[...slug]]/route.ts) — both are generateStaticParams work in a phase the build dies before.

Two of #12683's negative findings are confirmed by measurement below: --max-old-space-size and experimental.cpus do not bound this peak. experimental.cpus additionally cannot matter here for a structural reason — its static-generation workers have not spawned yet when the kill lands.

Experiment matrix (local cold builds, 403 MDX, peak single-process RSS)

config peak compile verdict
baseline 5191 MB 22.6s —
experimental.turbopackSourceMaps: false 4757 MB 19.5s ✅ -434 MB (-8.4%), no downside
turbopackScopeHoisting: false 4806 MB 19.1s ❌ no effect
NODE_OPTIONS=--max-old-space-size=2048 4734 MB 18.0s ❌ no effect (peak is Rust-side, confirming #12683)
turbopackFileSystemCacheForBuild: true 4779 MB 22.2s ❌ no effect
+ turbopackMinify: false 3785 MB 16.2s ⚠️ -972 MB more, but rejected
includeProcessedMarkdown: false — — ❌ build fails (getLLMText / /llms.txt depend on it)

Why turbopackMinify: false is rejected: client JS goes 5.8 MB -> 16 MB (+176%) — that is real user download weight. Restricting it to the server side (where the memory actually goes: 589 MB of server chunks vs 5.8 MB client) is not available: experimental.serverMinification is consumed only by dist/build/webpack-config.js and is never read on the Turbopack path.

Why sourcemaps are free here

The build emits 260 sourcemap files totalling 348 MB into .next/server. Production serverless functions have no use for them. Next's documented build-time default for turbopackSourceMaps follows productionBrowserSourceMaps (false), but the server-side maps are emitted regardless — setting the flag explicitly is what suppresses them.

Side benefits beyond the memory: 348 MB smaller build output, and compile time drops 22.6s -> 19.5s locally.

Proposed change

// apps/docs/next.config.mjs
experimental: {
  cpus: 2,
  turbopackSourceMaps: false,
}

Honest limitation: local measurement is macOS/arm64; Vercel is Linux/x86_64, and the absolute figures differ substantially (local 4.7 GB vs the 7.6 GB #12683 measured). The ratio is the transferable part, and even it is not a guarantee. Scaling -8.4% onto 7592 MB gives ~6950 MB against an 8192 MB container — roughly 15% headroom. That is thin. It is worth trying first because the cost of being wrong is one build cycle, and #12683's option A (bigger build machine) stays available and unaffected.

Also worth noting: Vercel takes 11 minutes to compile what takes 20s locally. A 4-vs-10 core gap does not explain 30x; the rest is thrashing near the container ceiling. Lowering the peak should cut build time substantially as a side effect.

Correction to #12698

#12698's action 2 ("fix the Ignored Build Step — it skips dependency-only changes") rests on a misreading. The project's actual rule is:

if [ "$VERCEL_ENV" == "production" ]; then exit 1; else exit 0; fi

exit 1 = build, exit 0 = skip. So production always builds and Preview always skips — it is not a path-scoped rule. The Canceled by Ignored Build Step that #12698 read off PR #12684's head was that PR's Preview deployment being skipped, which is the intended behaviour. Pushes to main do trigger production builds; they have simply been OOM-failing. No dashboard change is needed for the ignore step.

Re-check

curl -s https://objectstack.ai/ | grep -o 'data-dpl-id="[^"]*"'

Refs: #12683 (standing-risk card — this adds an option E its set did not contain) · #12698 (ignore-step correction) · #12677 · #12333 · epic #12243.

Activity

  1. self-assigned this
    on Aug 27, 2026
  2. os-zhuang commented on Aug 27, 2026

    @os-zhuang
    ContributorAuthor

    Claimed. session 3100934e-af9d-4104-9145-a37b8280a504 · branch claude/issue-12711-docs-build-sourcemaps · worktree sort-dot-path-fallback-fc72e9.

    Scope for this claim: the experimental.turbopackSourceMaps: false change to apps/docs/next.config.mjs plus its changeset — nothing else. The build-machine decision (#12683 option A) stays with the maintainer and is untouched by this PR; it remains available if the measured headroom turns out insufficient on Linux.

  3. os-zhuang commented on Aug 27, 2026

    @os-zhuang
    ContributorAuthor

    Result of the merged fix: the failure mode changed, the outage did not end

    PR #12712 merged as 18b6d89c. The production build it triggered (dpl_BF8NovndAJa5ezngzUJPphn41hXB) failed, but not the way its predecessors did:

    before the fix after the fix
    errorCode out_of_memory BUILD_EXCEEDED_MAXIMUM_TIME
    how it ended SIGKILL at 2m / 11m23s ran the full 46.1 min and timed out

    The flag is confirmed applied on Vercel — the build log prints ⨯ turbopackSourceMaps alongside · cpus: 2, byte-identical to the local run where source maps measurably went 260 files/348 MB → 0 and peak RSS dropped 434 MB.

    Reading

    This is the signature of a process that stopped crossing the OOM line but is still pressed against it. It entered Creating an optimized production build ... at 13:28:53 and emitted nothing for the next 27+ minutes; the same compile takes 20s locally. Reclaim pressure near the container ceiling is the only thing that accounts for a 100x slowdown on work that is not I/O bound, and it is exactly what the earlier 11m23s-of-silence-then-SIGKILL was already showing in weaker form.

    So the −8.4% was real and it was not enough: it bought just enough headroom to avoid the kill and not enough to let the compiler run. The ~15% projected margin this card flagged as "thin" turned out to be thin in the direction that matters.

    What this retires

    • All remaining free in-repo knobs. turbopackScopeHoisting, --max-old-space-size, turbopackFileSystemCacheForBuild measured as noise (matrix in the issue body); turbopackMinify: false would give another 972 MB but inflates client JS 5.8 MB → 16 MB (+176%) and cannot be confined to the server side; includeProcessedMarkdown: false breaks getLLMText. There is no further free number to lower.
    • The hypothesis that the fix alone could clear the ceiling. It cannot. Measured, not projected.

    What it confirms

    Memory is the binding constraint, and the container is undersized for this build — now demonstrated from both sides: below the line it is killed, at the line it is too slow to finish. #12683 option A (a larger build machine) is the remaining route, and this run is the evidence its decision was waiting on. Current machine: standard, 4 cores / 8192 MB, on a pro team.

    The merged change stays: it is a net win independently of the machine (348 MB less build output, −434 MB peak, faster compile), and on a larger container it keeps that margin rather than spending it.

    ⚠️ Note for whoever acts on this: the site is still serving dpl_2nfWjGjSwZjakVEmUG1kBWD6697r, unchanged since before the outage.

  4. hotlong commented on Aug 27, 2026

    @hotlong
    Contributor

    Resolved — the outage is over, and the fix was the machine

    The maintainer moved the docs project to an Enhanced build machine (8 vCPU / 16384 MB, project-level only; the other six projects stay on standard). Vercel had independently reached the same conclusion, recording buildMachineElasticReason: "build-timeout-failure" on the project.

    Result, measured:

    Standard (4 vCPU / 8 GB) Enhanced (8 vCPU / 16 GB)
    outcome out_of_memory, then BUILD_EXCEEDED_MAXIMUM_TIME READY
    wall time 46–50 min 3.7 min
    cost/build 184–200 CPU-min ≈ $0.64–0.70 (wasted) 30 CPU-min ≈ $0.105

    https://objectstack.ai/ now serves dpl_3YG2asYDmqBzoQt9eiicVLgZYivP — off the pinned dpl_2nfWjGjSwZjakVEmUG1kBWD6697r for the first time since the outage began. HTTP 200, /docs renders. The re-check in #12698 is satisfied.

    A larger machine turned out to be cheaper per build, not more expensive. Same $0.0035/CPU-minute across tiers, and doubling the cores cut wall time by ~13x because the slowness was reclaim pressure, not work. The earlier projection in this thread ($110–164/mo) was pessimistic for the same reason.

    What the evidence chain established

    1. The kill was always in the Turbopack compile phase, never the prerender — which retires the attribution in docs build: next build holds ~7.6 GB in ONE process — the docs deploy's next memory ceiling, and no in-repo knob bounds it #12683 and clears OG images / markdown routes as suspects.
    2. Peak scales as 2360 MB + 6.21 MB × pages locally, which extrapolates to ≈3682 MB + 9.7 MB × pages on Vercel and predicts 7591 MB at 403 pages against docs build: next build holds ~7.6 GB in ONE process — the docs deploy's next memory ceiling, and no in-repo knob bounds it #12683's independently measured 7592 MB. 403 pages was at 92.7% of an 8192 MB container — past the safe water line, which is why it had been fine for weeks and then was not.
    3. Free knobs are exhausted: turbopackSourceMaps: false (merged, −434 MB, kept — it now banks margin instead of spending it) was the only real one. turbopackScopeHoisting, --max-old-space-size, turbopackFileSystemCacheForBuild all measured as noise (±200 MB). turbopackMinify: false gives −972 MB but inflates client JS 5.8 → 16 MB. Splitting large MDX files does nothing — cost is per page (6.21 MB), not per byte; truncating the 292.9 KB v17.mdx moved the peak by +191 MB, i.e. noise.
    4. Turbopack's build-time filesystem cache raises peak memory (4591 → 6488 MB) while cutting compile 20s → 6.1s. It trades the resource we were short of for the one we were not; not usable here, though it may become attractive now that headroom exists.

    Capacity, going forward

    At ≈3682 MB + 9.7 MB/page against 16384 MB, the 75% water line sits near 870 pages; the site is at 403. That is the number to re-check against, and #12743 (avoiding unnecessary builds) reduces how often the ceiling is approached at all.

    Closing this card. Residual work is tracked in #12743.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions