Repository navigation
fix(runtime-host): cold-start artifact recovery sweeps O(all files) realpath/lstat, blocking Host readiness for minutes #4027
Description
Activity
Working on this.
Plan (no feature removal, no contract changes):
- Make authority-mode orphan discovery lazy: drop the eager full-tree scan in
recoverForWriteWithAuthority(findRecoverableOrphansUnlocked, which realpath+lstat+sha256-digests every unreferenced file) and reuse the existing per-createfindCompatibleRecoverableOrphansUnlockedpath, exactly like self-managed mode. Safety is preserved becauseadoptRecoverableOrphanUnlockedre-verifies lstat + digest at adoption time. - Parallelize the per-record
resolveArtifactRemovalEntry(realpath) loop inpreparePurgePathsUnlockedwith bounded concurrency, keeping the inode-aliasing integrity check intact.
Will open a PR.
- Make authority-mode orphan discovery lazy: drop the eager full-tree scan in
- added a commit that references this issue
on Aug 27, 2026 Follow-up profiling data: the residual cold-start cost is the O(M×N) metadata rewrite, not recovery-time hashing
With #4031 (revised head, bounded purge resolution) and #4033 applied locally, cold-starting a Runtime Host against the same real store (~11.7k artifact records, hundreds of accumulated Sessions) now reaches full TUI readiness in ~130 s (down from 10+ minutes originally, and from ~157 s before the recovery-dedupe fix). A windowed CPU profile of the Runtime Host during that period attributes the remaining cost almost entirely to the artifact metadata store:
- 0–90 s window: saturated by
artifact_recordstable churn. Roughly 34% of samples in native statementrun(the per-record INSERT), ~11% inexec(theDELETE FROM artifact_records), ~6.5% indecodeArtifactRecordJsons, ~5.4% in theSELECT record_json ... .all()read-back, and ~5.6% in sha256 — which isartifactIdentityKey()re-hashing every record id on every rewrite (plus recovery-time orphan digests). - 90–105 s: tapering off; after ~105 s the Host is essentially idle and the remaining wall time is client-side connection/rendering.
The mechanism, for the record: every artifact mutation (each retiring Session's sidecar purge, each recovery adoption) performs
reloadForMutationUnlocked()→readAll()(decode all N records) and thenreplaceAll()(DELETE+ re-INSERT all N records, each with a freshJSON.stringify+ sha256 identity key). With M serialized startup mutations and N ≈ 11.7k records, this is O(M×N) JSON decode/encode + SQL churn — hundreds of full-table rewrites of an ~11.7k-row table, which is exactly the ~90 s seen above.This confirms the follow-up direction already outlined in the #4031 review thread:
- Batched multi-Session purge via the existing multi-ID purge intent (one guard scan, one unlink plan, one metadata commit per retirement batch) removes the M multiplier on the filesystem-resolution side.
- Metadata change tracking (write back only the records a mutation actually touched, instead of
DELETE+ full re-INSERT) removes the M multiplier on the metadata side; each batched purge would then pay O(changed) instead of O(N). - The per-rewrite sha256 identity key (
artifactIdentityKey) is another small constant factor worth folding into the same change — the key is a pure function of the record id and could be computed once per record rather than once per record per rewrite.
(1)+(2) together should take the residual ~90 s metadata phase down to roughly the cost of one full-table pass. Happy to prototype either piece if the direction sounds right.
This comment was prepared with AI assistance (Kimi k3-256k), including profiling and analysis.
- 0–90 s window: saturated by
Filed the two contract-level follow-ups from the profiling above as separate design proposals, to keep this issue as the investigation thread: #4037 (change-tracked artifact metadata write-back — removes the O(N) full-table rewrite per mutation) and #4038 (batched multi-Session purge during retirement — removes the M multiplier). Direction questions for maintainers are in each issue.
Status note, tying the threads together. Discussion #4030 has converged on retiring the Artifact authority; per that discussion this issue stays open as the user-visible bug and becomes unreachable once the cutover lands. In the meantime the stopgaps are: #4031 (bounded purge resolution — approved at
7dad44778, CI green), which took the motivating workspace's cold start from 10+ minutes to roughly 2.5 minutes — the bulk of that improvement came from the pooled resolver, since the reverted digest commit accounted for only ~4 s here (most files are referenced and skip digesting) — and #4033 (recovery header listing without the duplicate decode), worth another ~20 s. The residual ~90 s is the metadata full-table rewrite storm profiled above; per #4030's cutover invariants that cost is expected to be removed by the retirement migration rather than patched in place.This comment was prepared with AI assistance (Kimi k3-256k).
- added a commit that references this issue
on Aug 28, 2026 - addedbugSomething isn't workingSomething isn't workinghelp wantedExtra attention is neededExtra attention is needed
on Aug 29, 2026 - addedstaleNo qualifying activity within the lifecycle policy windowNo qualifying activity within the lifecycle policy window
on Sep 27, 2026 This issue has had no human activity for 30 days and has been marked stale. It will be closed in 30 days unless someone comments.
If the issue is still current, please confirm it against the latest
mainand add any information that would help move it forward. Assigned issues and issues labelledpinnedare exempt from this policy.take
- removedstaleNo qualifying activity within the lifecycle policy windowNo qualifying activity within the lifecycle policy window
on Oct 9, 2026
What happened
Cold-starting a Runtime Host against a workspace with a large accumulated artifact store blocks Host readiness for ~10 minutes at 60-70% CPU. Connected TUI clients render nothing during this window (
maka --resume <id>appears to hang).Observed on a long-lived workspace with ~11.7k artifact files (~6 GB on disk, comparable
artifact_recordsrow count). A long-running Host amortizes these costs in memory; when the Host is replaced (for example a version upgrade with a compatibility-epoch cutover), the new Host pays the full cost synchronously during startup recovery.CPU profile of the Host main thread during the stall (CDP
Profiler, 500µs sampling):preparePurgePathsUnlocked(packages/storage/dist/artifact-store.js:528)resolveArtifactRemovalEntry→realpath/lstatper record (node:internal/fs/promises)decodeArtifactRecordJsons(artifact-metadata-codec.js) — full in-memory decode of all ~11.7k recordsreadEventsForRecovery/readSqliteAgentRunEventsreplaying 56kcore_agent_run_events+ 29ksession_messagesbefore servingRoot causes in
packages/storage/src/artifact-store.ts:hasCanonicalRecoveryResidueUnlocked,recoverMetadataTempsUnlocked, publication recovery)readdirs every session directory and runsrealpath+lstaton every artifact file — O(files) syscalls on every Host start.preparePurgePathsUnlockedperforms its referential-integrity check by resolvingresolveArtifactRemovalEntry(realpath) for all records, not just the purge set — O(all records) perpurge()call, so retiring M sessions costs M × N realpaths.metadataRepository.readAll()).Severity scales with artifact count. Fresh/small installs are unaffected (sub-millisecond sweeps), but heavy long-lived workspaces degrade on every Host restart (version update, epoch cutover, crash recovery), and the TUI gives no progress indication while it waits.
How to reproduce
maka --resume <session-id>.Profiling one-liner used:
Environment
4cc781f31(main, 2026-08-27)Logs, screenshots, or additional context
Top self-time frames from the profile (5s window during the stall):
(libuv threadpool fs syscalls account for additional CPU not visible to the main-thread inspector.)
Suggested directions:
realpath/lstatsweeps.preparePurgePathsUnlocked, restrict referential-integrity checks to records sharing the purge set's relative paths (metadata query) rather than resolving every record.