Repository navigation
docs deploy: Vercel production build OOM-killed — @objectstack/spec:build exits 137, every production deploy since 2026-08-25 ~16:55 fails (root cause behind #12333) #12677
Copy link
Copy link
Closed
Labels
documentationImprovements or additions to documentationImprovements or additions to documentationdomain:devxpm:dispatchedpriority:p1High: required for production / M2High: required for production / M2repo:objectstack
Description
Activity
- addeddocumentationImprovements or additions to documentationImprovements or additions to documentation
on Aug 27, 2026 - addedpriority:p1High: required for production / M2High: required for production / M2
on Aug 27, 2026 Claiming this issue.
- Session:
session_01DKWDdUJ2XNRESVVWUvcpnh - Branch:
claude/issue-12677-docs-build-oom
Starting with the versioned build command in
apps/docs/vercel.jsonand a measured baseline ofpnpm turbo run build --filter=@objectstack/docsunder a memory observer.Generated by Claude Code
Generated by Claude Code
- Session:
os-dev-report
{ "issue": 12677, "status": "done", "branch": "claude/issue-12677-docs-build-oom", "pr": "https://github.com/objectstack-ai/objectstack/pull/12684", "premise_still_valid": true, "summary": "Root cause confirmed and fixed at the process the Vercel log names: packages/spec ran its DTS pass under NODE_OPTIONS=--max-old-space-size=12288 -- a 12 GB heap ceiling inside an ~8 GB build container. A ceiling above the container's memory does not permit a bigger build; V8 defers major GCs below its ceiling, so the kernel SIGKILLs the process before V8 reaches it, which is exactly exit 137 with no diagnostic. Lowering it to 6144 cuts that phase's peak from 6965 MB to 5654 MB (-19%) with byte-identical declaration output and no wall-time cost, and makes overflow fail loud (ERR_WORKER_OUT_OF_MEMORY) instead of silently. Lever 1 (turbo concurrency) was RULED OUT by measurement, not skipped: the docs build graph is 2 tasks in a linear chain (spec is the docs app's only workspace dependency), so the killed process was alone in the container and --concurrency would have changed nothing; the build command is versioned in apps/docs/vercel.json, so no dashboard-only configuration was involved in the command itself. Two process notes: the issue arrived with assignees EMPTY (PM-side half-state, worked as instructed), and the dispatch prompt's 'assign + claim comment' instruction conflicts with my standing contract that the assignee field belongs to the PM -- I resolved it per the standing contract (claim comment with session ID and branch posted, assignee not written) and flag the conflict here rather than choosing silently.", "tests": "All measurements taken at commit a0a64684 on a clean tree. Harness validated in BOTH directions before any number was trusted -- positive control 'MEASURE label=selftest_ok limit_mb=8192 exit=0 peak_rss_mb=589 oom_kill=0', negative control 'MEASURE label=selftest_oom limit_mb=1024 exit=137 peak_rss_mb=1021 oom_kill=1' (reproduces the Vercel signature). Container emulation is two halves -- cgroup-v1 cap AND a private mount namespace whose /proc/meminfo reports the container size (verified active: 'seen totalmem GB 8.00'); a cap alone bounds what a process may use without changing what it believes it has, which matters because the docs phase grows until killed rather than sizing down. BEFORE/AFTER, faithful 8192 MB container: 'MEASURE label=final_dts_before container_mb=8192 exit=0 seconds=126 peak_rss_mb=6965 oom_kill=0' vs 'MEASURE label=final_dts_after container_mb=8192 exit=0 seconds=124 peak_rss_mb=5654 oom_kill=0'. Cap sweep choosing 6144 (cgroup-only): 12288->7290 MB, 6144->5794 MB, 5120->5328 MB, 4096->exit 1 ERR_WORKER_OUT_OF_MEMORY. Output identity: declarations deleted before each run (the DTS pass runs clean:false), then 'DTSHASH cap=6144 decl_files=122 tree_sha256_16=37cf1007189f945c' and 'DTSHASH cap=5120 ... 37cf1007189f945c' -- byte-identical; check-dts-emitted reports 34/34 declared declaration files present. Full pipeline 'pnpm turbo run build --filter=@objectstack/docs --force' from cleaned dist/.next with the fix: 'Tasks: 2 successful, 2 total / Time: 5m02s', 'MEASURE label=final_pipeline_10g container_mb=10240 exit=0 peak_rss_mb=8135 oom_kill=0'. Gates: 30 derived families via 'node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack' -- all PASS rc=0, plus check:nul-bytes PASS and '@objectstack/spec typecheck' PASS. Exit codes captured before any pipe. TWO NON-MEASUREMENTS, recorded as NOT MEASURED rather than red: check-dev-prereqs rc=1 prints 'The workspace is not built -- 1 unmet precondition' (this worktree built only spec and docs, not all 67 packages; CI checks out fresh), and scripts/pm/check-half-states.mjs rc=3 prints 'Nothing was swept ... it is no reading at all' for lack of a GitHub token. Declared narrowing: repository-wide lint left to CI. All heavy runs serialized through scripts/pm/os-verify-lock.sh (VERDICT command-exit 0 on every hold).", "mcp_calls": "6 — issue_read get, issue_read get_comments, add_issue_comment (claim), issue_write (create #12683), create_pull_request, add_issue_comment (this report). Issue de-duplication and the PR body read-back went through repo-scoped REST, which this seat can read; note /search/ is BLOCKED for this session ('sessions are bound to their configured repositories'), so dedup used the repo-scoped list endpoint plus local grep, with a control term that hit.", "open_questions": [ { "question": "MAINTAINER-ONLY DASHBOARD STEP: the docs pipeline's peak is now owned by `next build`, which holds ~7.6 GB in ONE process and is bounded by nothing in this repo. Its demand is fixed, not adaptive (killed at 8192/6144/4096 MB rather than sizing down). What should back it off?", "options": [ "A — Raise the Vercel build container (dashboard action; NOT versioned in apps/docs/vercel.json, so no seat can do it in-repo). Buys headroom, changes no code, costs money.", "B — Build the docs app with webpack instead of Turbopack. Different memory profile, but a bundler change for a 400-page site, and next.config.mjs records that Vercel and CI both run Turbopack today.", "C — Shard the prerender so no single process holds the whole page set. Largest change; the only one that scales with page count." ], "recommendation": "A short-term, C long-term. This is NOT blocking the outage fix: that phase has been building successfully on Vercel at this size for weeks, so it is a standing risk, not a live failure -- but the margin is thin, the page count grows, and when it crosses it produces the identical exit-137-with-no-diagnostic signature with no in-repo number left to lower. Filed with all measurements as #12683 rather than ridden into this PR, because A is a maintainer decision and B/C are architecture decisions." } ], "out_of_scope_findings": [ "filed as #12683: `next build` holds ~7.6 GB in one process (turbopack workers are only 100-430 MB each); neither --max-old-space-size (Turbopack is Rust, outside V8) nor experimental.cpus bounds it, and Next 16.3.1 exposes no memoryLimit config option — the docs deploy's next memory ceiling" ] }Generated by Claude Code
Generated by Claude Code
Metadata
Metadata
Assignees
Labels
documentationImprovements or additions to documentationImprovements or additions to documentationdomain:devxpm:dispatchedpriority:p1High: required for production / M2High: required for production / M2repo:objectstack
What breaks, for whom
objectstack.aiproduction is pinned to a 2026-08-25 build: every Vercel production build since ~16:55 that day has died, so 8+ merged docs PRs from epic #12243 are invisible to users and crawlers (full measurement in #12333, which closes into this card). The maintainer read the Vercel build log on 2026-08-27 and pasted it verbatim in the PM session:Exit 137 = 128 + SIGKILL: the build container's OOM killer, not a compile error. The Vercel build command runs
turbo run build --filter=@objectstack/docs, which builds the docs app's whole dependency graph in parallel inside one fixed-memory container.@objectstack/spec:buildis the process the kernel chose to kill — which does not necessarily make it the process that grew.Task
Make the docs production build fit the container's memory, by changes inside the repo only. Candidate levers, in preferred order — pick what measurement supports, do not apply all blindly:
vercel.jsonat root or underapps/docs; if it turns out to be dashboard-only configuration, stop and report back instead of working around it).--concurrency=1(or2) trades minutes of build time for a bounded memory peak./usr/bin/time -v, or a cgroup cap emulating the container) at currentmain. If one package's build dominates —@objectstack/spec's dts generation is the usual suspect — bound it (split the dts step, tsup/tsc settings, an explicitNODE_OPTIONSheap cap so it fails loud instead of ballooning).mainmay identify exactly what grew.Acceptance
pnpm turbo run build --filter=@objectstack/docsfrom a clean checkout completes with a measured peak memory stated in the PR (before/after numbers), fitting the container envelope with headroom.content/docs/releases/; no new runtime dependencies; anything requiring a Vercel dashboard change is a report-back to the PM, not a workaround hacked into code.Verification after merge (inherited from #12333)
Once the fix merges and Vercel rebuilds:
must show an id other than
dpl_2nfWjGjSwZjakVEmUG1kBWD6697r. That id moving is the definition of done for the outage; the epic #12243 site-level verification resumes from there.Refs: #12333 (measurement card, closes into this one) · epic #12243.