Skip to content

fix(runtime-host): cold-start artifact recovery sweeps O(all files) realpath/lstat, blocking Host readiness for minutes #4027

Description

@me2seeks

What happened

Cold-starting a Runtime Host against a workspace with a large accumulated artifact store blocks Host readiness for ~10 minutes at 60-70% CPU. Connected TUI clients render nothing during this window (maka --resume <id> appears to hang).

Observed on a long-lived workspace with ~11.7k artifact files (~6 GB on disk, comparable artifact_records row count). A long-running Host amortizes these costs in memory; when the Host is replaced (for example a version upgrade with a compatibility-epoch cutover), the new Host pays the full cost synchronously during startup recovery.

CPU profile of the Host main thread during the stall (CDP Profiler, 500µs sampling):

  • preparePurgePathsUnlocked (packages/storage/dist/artifact-store.js:528)
  • resolveArtifactRemovalEntry → realpath / lstat per record (node:internal/fs/promises)
  • decodeArtifactRecordJsons (artifact-metadata-codec.js) — full in-memory decode of all ~11.7k records
  • Followed by readEventsForRecovery / readSqliteAgentRunEvents replaying 56k core_agent_run_events + 29k session_messages before serving

Root causes in packages/storage/src/artifact-store.ts:

  1. Startup recovery (hasCanonicalRecoveryResidueUnlocked, recoverMetadataTempsUnlocked, publication recovery) readdirs every session directory and runs realpath + lstat on every artifact file — O(files) syscalls on every Host start.
  2. preparePurgePathsUnlocked performs its referential-integrity check by resolving resolveArtifactRemovalEntry (realpath) for all records, not just the purge set — O(all records) per purge() call, so retiring M sessions costs M × N realpaths.
  3. The artifact store eagerly decodes all metadata records into memory at open (metadataRepository.readAll()).

Severity scales with artifact count. Fresh/small installs are unaffected (sub-millisecond sweeps), but heavy long-lived workspaces degrade on every Host restart (version update, epoch cutover, crash recovery), and the TUI gives no progress indication while it waits.

How to reproduce

  1. Accumulate a workspace artifact store with ~10k files (long-lived heavy use: many sessions, side conversations, tool outputs).
  2. Stop the running Runtime Host (or trigger an epoch cutover via a version upgrade).
  3. Run maka --resume <session-id>.
  4. Observe: new Host spins at 60-70% CPU for ~10 minutes; TUI stays blank until Host recovery completes.

Profiling one-liner used:

kill -USR1 <host-pid>   # enable inspector on 127.0.0.1:9229
# then CDP Profiler.start/stop over ws://127.0.0.1:9229/<id>

Environment

  • Commit: 4cc781f31 (main, 2026-08-27)
  • Node: 26.3.0
  • OS: Linux
  • Surface: Runtime Host / TUI

Logs, screenshots, or additional context

Top self-time frames from the profile (5s window during the stall):

12.1%  run
 5.2%  run
 1.3%  realpath
 1.1%  lstat
 1.0%  lstat                    node:internal/fs/promises:1670
 0.9%  realpath                 node:internal/fs/promises:1826
 0.8%  preparePurgePathsUnlocked  packages/storage/dist/artifact-store.js:528
 0.6%  decodeArtifactRecordJsons  packages/storage/dist/artifact-metadata-codec.js:40

(libuv threadpool fs syscalls account for additional CPU not visible to the main-thread inspector.)

Suggested directions:

  • Make startup recovery lazy or index-driven (metadata in SQLite) instead of per-file realpath/lstat sweeps.
  • In preparePurgePathsUnlocked, restrict referential-integrity checks to records sharing the purge set's relative paths (metadata query) rather than resolving every record.
  • Surface Host bootstrap progress to waiting TUI clients so a cold start does not look like a hang.

Activity

  1. me2seeks commented on Aug 27, 2026

    @me2seeks
    ContributorAuthor

    Working on this.

    Plan (no feature removal, no contract changes):

    1. Make authority-mode orphan discovery lazy: drop the eager full-tree scan in recoverForWriteWithAuthority (findRecoverableOrphansUnlocked, which realpath+lstat+sha256-digests every unreferenced file) and reuse the existing per-create findCompatibleRecoverableOrphansUnlocked path, exactly like self-managed mode. Safety is preserved because adoptRecoverableOrphanUnlocked re-verifies lstat + digest at adoption time.
    2. Parallelize the per-record resolveArtifactRemovalEntry (realpath) loop in preparePurgePathsUnlocked with bounded concurrency, keeping the inode-aliasing integrity check intact.

    Will open a PR.

  2. me2seeks commented on Aug 27, 2026

    @me2seeks
    ContributorAuthor

    Follow-up profiling data: the residual cold-start cost is the O(M×N) metadata rewrite, not recovery-time hashing

    With #4031 (revised head, bounded purge resolution) and #4033 applied locally, cold-starting a Runtime Host against the same real store (~11.7k artifact records, hundreds of accumulated Sessions) now reaches full TUI readiness in ~130 s (down from 10+ minutes originally, and from ~157 s before the recovery-dedupe fix). A windowed CPU profile of the Runtime Host during that period attributes the remaining cost almost entirely to the artifact metadata store:

    • 0–90 s window: saturated by artifact_records table churn. Roughly 34% of samples in native statement run (the per-record INSERT), ~11% in exec (the DELETE FROM artifact_records), ~6.5% in decodeArtifactRecordJsons, ~5.4% in the SELECT record_json ... .all() read-back, and ~5.6% in sha256 — which is artifactIdentityKey() re-hashing every record id on every rewrite (plus recovery-time orphan digests).
    • 90–105 s: tapering off; after ~105 s the Host is essentially idle and the remaining wall time is client-side connection/rendering.

    The mechanism, for the record: every artifact mutation (each retiring Session's sidecar purge, each recovery adoption) performs reloadForMutationUnlocked() → readAll() (decode all N records) and then replaceAll() (DELETE + re-INSERT all N records, each with a fresh JSON.stringify + sha256 identity key). With M serialized startup mutations and N ≈ 11.7k records, this is O(M×N) JSON decode/encode + SQL churn — hundreds of full-table rewrites of an ~11.7k-row table, which is exactly the ~90 s seen above.

    This confirms the follow-up direction already outlined in the #4031 review thread:

    1. Batched multi-Session purge via the existing multi-ID purge intent (one guard scan, one unlink plan, one metadata commit per retirement batch) removes the M multiplier on the filesystem-resolution side.
    2. Metadata change tracking (write back only the records a mutation actually touched, instead of DELETE + full re-INSERT) removes the M multiplier on the metadata side; each batched purge would then pay O(changed) instead of O(N).
    3. The per-rewrite sha256 identity key (artifactIdentityKey) is another small constant factor worth folding into the same change — the key is a pure function of the record id and could be computed once per record rather than once per record per rewrite.

    (1)+(2) together should take the residual ~90 s metadata phase down to roughly the cost of one full-table pass. Happy to prototype either piece if the direction sounds right.

    This comment was prepared with AI assistance (Kimi k3-256k), including profiling and analysis.

  3. me2seeks commented on Aug 27, 2026

    @me2seeks
    ContributorAuthor

    Filed the two contract-level follow-ups from the profiling above as separate design proposals, to keep this issue as the investigation thread: #4037 (change-tracked artifact metadata write-back — removes the O(N) full-table rewrite per mutation) and #4038 (batched multi-Session purge during retirement — removes the M multiplier). Direction questions for maintainers are in each issue.

  4. me2seeks commented on Aug 28, 2026

    @me2seeks
    ContributorAuthor

    Status note, tying the threads together. Discussion #4030 has converged on retiring the Artifact authority; per that discussion this issue stays open as the user-visible bug and becomes unreachable once the cutover lands. In the meantime the stopgaps are: #4031 (bounded purge resolution — approved at 7dad44778, CI green), which took the motivating workspace's cold start from 10+ minutes to roughly 2.5 minutes — the bulk of that improvement came from the pooled resolver, since the reverted digest commit accounted for only ~4 s here (most files are referenced and skip digesting) — and #4033 (recovery header listing without the duplicate decode), worth another ~20 s. The residual ~90 s is the metadata full-table rewrite storm profiled above; per #4030's cutover invariants that cost is expected to be removed by the retirement migration rather than patched in place.

    This comment was prepared with AI assistance (Kimi k3-256k).

  5. added theissue type on Aug 29, 2026
  6. added
    staleNo qualifying activity within the lifecycle policy window
    on Sep 27, 2026
  7. github-actions commented on Sep 27, 2026

    @github-actions

    This issue has had no human activity for 30 days and has been marked stale. It will be closed in 30 days unless someone comments.

    If the issue is still current, please confirm it against the latest main and add any information that would help move it forward. Assigned issues and issues labelled pinned are exempt from this policy.

  8. BigDataDZ commented on Oct 9, 2026

    @BigDataDZ
    Contributor

    take

  9. removed
    staleNo qualifying activity within the lifecycle policy window
    on Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't workinghelp wantedExtra attention is needed

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions