Repository navigation
feat(storage): report Host storage usage and per-task sizes - #5832
Conversation
Add a read-only Runtime Host query, `storage.usage.query`, that reports how much space the State Root uses, and show it in Desktop Settings. Nothing here deletes, compacts or changes stored data. Storage measures cheaply without write-path counters or triggers: file sizes of runtime.sqlite (+wal/shm) and memory.sqlite, `octet_length`/`length` sums over the per-Session tables through their existing session_id indexes, artifact `sizeBytes`, context-offload usage, `PRAGMA freelist_count * page_size` for reclaimable space, and a count of subagent worktree directories. The totals are non-overlapping: transcript, runtime and usage history are logical bytes inside runtime.sqlite and `database` is the remainder of that file set. The Host shares one totals scan between concurrent callers and for 15s; per-Session figures are always read fresh. Input takes up to 100 sessionIds and the result echoes them in order. The query is admitted for remote owners like `host.resources.query`. Compatibility epoch 198 -> 199 for the new operation. Desktop adds a storage-usage feature slice: a Storage group on the Data page after the data location group, and a per-task size on archived task rows, measured only for rows that mount and batched into one query per Host. Copy notes that per-task offloaded context is logical and that usage history is kept after deletion. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review of the first M1 cut showed the full-table payload sums behind the transcript/runtime/usage_history totals read every leaf page in one synchronous call and blocked the Host event loop. - Totals report only facts that are cheap to learn: the runtime.sqlite file set (exact), artifact metadata sizes (not exact), context-offload blob bytes plus its SQLite file set, and the memory file set. The -wal/-shm/-journal sidecar list is shared with the long-term memory store. Reclaimable space and the worktree count are unchanged. - Split the protocol into `storage.usage.query` (totals) and `storage.usage.sessions.query` (per-Session), so measuring rows never computes totals. The 15s totals cache is gone; concurrent requests still share one scan, so Refresh re-measures. - Per-Session measurement yields after every statement, is capped at 25 Sessions per request, and omits Sessions the Host does not hold instead of reporting 0 B. The result is an ordered subsequence of the request. - Desktop measures rows one bounded request at a time, caches results for 60s and failures for a 30s cooldown, and a failing Host no longer hides other Hosts' sizes. The loader lives in the feature provider, so the Tasks page only gains the size label. - Reuse and extend @maka/ui formatBytes (GB/TB, optional locale) instead of a second formatter; kind order and copy are exhaustive over STORAGE_USAGE_KINDS; the Storage group drops the previous Host's figures when the Host or its generation changes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
hqhq1025
left a comment
There was a problem hiding this comment.
Reviewed the storage-footprint reader, Host protocol/coordinator, Desktop routing and presentation at bc635bc. The read-only split and bounded per-session requests are useful, but I found one issue in the total footprint. This is not an approval to merge.
The focused storage/Host/Desktop tests pass (12/12), as do build:test, ASF headers, app-shell hooks, focused Biome, diff-check, and a merge-tree against current main. The current-head hosted test check is red in an ACP Plan paging test outside the changed files; its cause is not established. I did not exercise a multi-GB workspace or packaged Desktop.
Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.
| { | ||
| kind: 'context_offload' as const, | ||
| // Inline blobs are counted in both terms, so this can overstate. | ||
| bytes: contextUsage.physicalBytes + contextFileBytes, |
There was a problem hiding this comment.
[P2] Do not add inline context bytes to the SQLite file size. Normal tool_result_archive blobs use inline storage (sqlite-context-offload-store.ts:1389-1390), and their bytes are already in context-offload.sqlite/WAL while also counted in context_store_usage.physical_bytes (:625-638). Adding that counter to the SQLite file set double-counts every inline blob, so the Storage page's "Total" overstates disk usage, not merely by rounding. A real 4,000,000-byte inline blob produced physicalBytes=4,000,000 and context_offload.bytes=12,265,376; the measured SQLite file set was 8,265,376 bytes. Please count only managed external file bytes in addition to the SQLite file set (or otherwise remove the overlap), and cover an inline blob in the footprint test.
Astro-Han
left a comment
There was a problem hiding this comment.
Reviewed exact head bc635bc492d9dc471f33bff9607324e4c1951661. This adds to the existing inline-BLOB double-count P2.
Trust boundary: looks sound. No client input becomes a filesystem path; the Host measures only fixed names inside its own State Root. Session ids are validated, capped at 25 per request and bound as SQL parameters. Remote owners can call the new operations and collaboration guests cannot. Worktrees are counted, never walked. There is no migration, and the per-task queries use session_id-leading indexes.
P2 (performance, inline): each per-task query sums up to 25 Sessions in one synchronous SQLite statement. The archived-tasks page measures every archived task, so it can block the Host event loop for about 1.3 s per statement on large tasks (reproduced locally with synthetic data). I haven't checked real-world task sizes.
P3:
- Epoch 198 → 199 collides with other open PRs.
- The renderer's size cache never evicts.
- The tests insert no
runtime_partial_snapshots/core_agent_run_eventsrows, so a wrong table or column name there would still pass.
Checks run:
- Pass: the ASF header check,
git diff --check, the app-shell hooks and renderer-architecture checks, locale hygiene, the desktop typecheck, and the protocol epoch check. - Up to date: the Windows test inventory.
- Tests: new 13/13, related runtime-host/desktop 139/139, storage 71/71.
- Not run: the full
npm test.
Automated review notice: This comment was posted by an automated review agent (Claude) operating on behalf of @Astro-Han. It is not an independent human review and does not replace one.
| ): Promise<Map<string, number>> { | ||
| const totals = new Map<string, number>(); | ||
| for (const source of sources) { | ||
| const rows = this.lease.database |
There was a problem hiding this comment.
P2: each statement sums payload bytes for up to 25 Sessions in one synchronous .all(). octet_length only skips overflow pages, so the cost grows with the byte volume of those Sessions' rows. Locally, 25 Sessions × 20k runtime events × ~1.5 KB blocked the Host event loop for about 1.3 s in a single statement (~67 ms at 4k events). The archived-tasks settings page then requests every archived task in 25-id pages, which in effect streams whole tables through the Host. Could this measure one Session per statement with a yield in between, or use a smaller bounded chunk? The UI could also measure only visible rows, or wait for a user action. A test asserting that the Host still yields during a multi-Session measurement would lock this in.
| // Increment when the same protocol version no longer guarantees safe Client-Host | ||
| // interoperability. Mismatches are rejected before domain commands are admitted. | ||
| export const RUNTIME_HOST_COMPATIBILITY_EPOCH = 198 as const; | ||
| export const RUNTIME_HOST_COMPATIBILITY_EPOCH = 199 as const; |
There was a problem hiding this comment.
P3 (coordination): several other open PRs also claim epoch 199/200, so whichever merges second has to renumber. Both new operations are additive; protocol-compatible-changes/recall-query-operation.json shows the exemption route for that case. Keeping the bump is fine, but it will need re-pinning if another PR lands first.
| options: { readonly now?: () => number } = {}, | ||
| ): SessionStorageLoader { | ||
| const now = options.now ?? Date.now; | ||
| const entries = new Map<string, Entry>(); |
There was a problem hiding this comment.
P3 (nit): nothing ever deletes from entries, so it grows with every task shown for the lifetime of the window. Consider dropping expired entries.
- context_offload now reports the context-offload SQLite file set plus the managed value files beside it. Inline blobs already live in the SQLite files, so adding the Store's physical-byte counter double counted them. - Per-Session measurement runs one Session per statement and yields to the event loop after each, so a large task no longer holds the Host for a whole batch. - The archived-task list measures a row only once it scrolls into view (IntersectionObserver); visible rows still share bounded requests. - The renderer size cache evicts expired entries on each flush and is capped at 1,000 entries. - A Host whose storage query fails is skipped for a cooldown. An older Host rejects the unknown operation by closing the connection, the same way it treats recall.query, so retrying on every visible row would keep dropping it; the UI shows its unavailable state instead. - Drop the compatibility epoch bump. The two operations are an additive extension, declared in protocol-compatible-changes at epoch 198. - Footprint tests now cover runtime_partial_snapshots and core_agent_run_events rows, inline versus managed context blobs, and that multi-Session measurement yields between statements. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Thanks @hqhq1025 and @Astro-Han, both P2s were real. Everything below is addressed in P2: inline context blobs counted twice (@hqhq1025). P2: per-task measurement blocking the Host (@Astro-Han).
P3s
CI note: the earlier red
The rerun is green. Checks on this head:
|
hqhq1025
left a comment
There was a problem hiding this comment.
This revision fixes the duplicate counting of inline context blobs by adding only the SQLite file set and managed value files. It also splits per-session SQL aggregation into yielding statements, adds visibility-gated task measurement, and bounds the renderer cache. The focused storage/UI tests, local build, and current-head hosted test pass. I found one remaining compatibility issue below; this is not a merge approval. I did not measure a multi-GB session or exercise an older remote Host end to end.
Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.
| "packages/runtime-host/src/protocol/storage-usage.ts", | ||
| "packages/runtime-host/src/protocol/operations.ts" | ||
| ], | ||
| "reason": "Adds the storage.usage.query and storage.usage.sessions.query operation specs and registers their keys, including the remote-owner grant list. Two new operation keys are an additive extension: a peer that predates them never sends either operation and never receives a response for one, so every existing request and response decodes byte-for-byte as before. Nothing already on the wire changes shape or meaning." |
There was a problem hiding this comment.
P2: The compatibility rationale covers an old client talking to a new Host, but not a new Desktop talking to an old Host at the same epoch 198. Current main advertises epoch 198 without either storage operation; opening Data Settings sends storage.usage.query unconditionally (storage-usage-section.tsx:55-71, preload.ts:3716). That Host rejects the unknown key in requireOperationKey, and connection-session.ts closes the live connection. Visible task rows also send storage.usage.sessions.query; their 60-second failure cooldown only limits repeat disconnects, while Settings has no such guard. Please negotiate support before sending these optional queries, or make the pair fail compatibility admission rather than dropping a live connection.
Astro-Han
left a comment
There was a problem hiding this comment.
Re-reviewed exact head 5af33e59d9004d1de9054fa4a197a1877e453a62 against my earlier review on bc635bc4.
Resolved
- Event-loop blocking (my earlier P2). Each Session now gets its own statements, with a yield between them. With my earlier synthetic load (25 Sessions × 20k events × ~1.5 KB), the longest single stall dropped from ~1.28 s to ~132 ms. What remains grows only with one Session's rows: ~245 ms for a 300 MB Session. The code documents that, and exact per-Session sums can't avoid it. Archived-task sizes are now measured only when rows scroll into view, and a test covers this.
- Inline-blob double count. The SQLite file set plus the
context-offload-values/files are now counted once each. The new real-store test checks this. - Earlier P3s. The renderer cache now expires entries and is capped at 1,000. The tests now insert
runtime_partial_snapshotsandcore_agent_run_eventsrows. The epoch stays at 198, with a compatible-change declaration.
Old-Host compatibility. This adds to the P2 already raised on this head. storage-usage-section.tsx:73 re-queries on every mount and every Retry with no backoff. So a new Desktop connected to an older Host at epoch 198 drops that Host's connection every time Data Settings is opened or retried. The per-task path at least waits 60 s after a failure. The reason in storage-usage-operations.json:7 covers only an older Client, not a new Client talking to an older Host. Either capability negotiation, or keeping the epoch bump, would close this.
Nit (inline): the 200 px prefetch rootMargin in task-storage-size.tsx:61 has no effect if the list scrolls inside a nested container, because the observer's root is the viewport.
Checks run
check:asf-headers,git diff --check, the protocol epoch check, app-shell-hooks, renderer-architecture, locale hygiene and the Desktop typecheck all pass.- The Windows test inventory is current.
- Targeted tests pass: storage 28/28, runtime-host 147/147, desktop 7/7.
- Not run: the full
npm test, or anything against a real older Host.
Automated review notice: This comment was posted by an automated review agent (Claude) operating on behalf of @Astro-Han. It is not an independent human review and does not replace one.
| : {}), | ||
| }, | ||
| })); | ||
| services.loadUsage(host).then( |
There was a problem hiding this comment.
This issues storage.usage.query on every mount and every Retry, with no backoff or capability check. An older Host at epoch 198 doesn't know the operation, so it rejects it in requireOperationKey and closes the connection each time Data Settings is opened or retried. The per-task path at least has a 60 s failure cooldown. Could this be gated on the Host advertising the operation, or back off after the first rejection?
| "packages/runtime-host/src/protocol/storage-usage.ts", | ||
| "packages/runtime-host/src/protocol/operations.ts" | ||
| ], | ||
| "reason": "Adds the storage.usage.query and storage.usage.sessions.query operation specs and registers their keys, including the remote-owner grant list. Two new operation keys are an additive extension: a peer that predates them never sends either operation and never receives a response for one, so every existing request and response decodes byte-for-byte as before. Nothing already on the wire changes shape or meaning." |
There was a problem hiding this comment.
This reason holds for an older Client, which never sends these keys. It doesn't cover a new Client talking to an older Host at epoch 198: that Host rejects the unknown key and drops the connection (see the comment on storage-usage-section.tsx:73).
| setVisible(true); | ||
| } | ||
| }, | ||
| { rootMargin: VISIBILITY_MARGIN }, |
There was a problem hiding this comment.
Nit: with the default viewport root, this 200 px rootMargin has no effect when the list scrolls inside a nested container. Rows are still measured when they become visible, just without the prefetch.
Treating the two storage operations as a compatible extension was wrong for a new Desktop talking to an epoch-198 Host: the Host rejects the unknown key in `requireOperationKey` and closes the live connection, and Data Settings sends `storage.usage.query` on every open and retry. Restore the epoch bump (198 -> 199), so the mismatched pair is refused at the handshake, and drop the compatible-change declaration. Observe task rows against their nearest scrolling ancestor, so the 200px prefetch margin applies inside the Settings scroll container. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Thanks @hqhq1025 and @Astro-Han, you're right: the compatible-extension declaration only covered an old Client talking to a new Host. The protocol has no way for a Host to advertise optional operations, and adding one would be new surface for a read-only feature. So
Re-ran on this head:
|
# Conflicts: # docs/astryx-surface-file-inventory.md
hqhq1025
left a comment
There was a problem hiding this comment.
Reviewed commit 95cc9be. This merge brings current main into the storage-usage branch and resolves the Astryx inventory conflict. The prior compatibility finding is addressed: the new storage operations use compatibility epoch 199, while epoch-198 Hosts are rejected during handshake before domain commands are admitted (protocol/index.ts:107-110, client/connection.ts:1372-1395, server/host-kernel.ts:553-565). The lazy task-size observer also now roots its prefetch margin at the nearest scrolling ancestor. I found no new P0–P3 issue on this head.
Clean Node 24 installation and build:test, 18 focused storage/handshake/UI tests, Astryx inventory coverage, Biome, diff-check, and a merge-tree against current main pass locally. Hosted Windows recovery passes; hosted test was still running at review time, so merge readiness remains contingent on that result. I did not run a real epoch-198 remote Host, a packaged Desktop, or a multi-GB storage workload.
Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.
Astro-Han
left a comment
There was a problem hiding this comment.
Re-reviewed exact head 95cc9be827d3edc3e8b316d742aa9fa2266ac488 (incremental from 5af33e59). No P0–P3 findings.
- Old-Host disconnect resolved. The unknown-operation rejection only happens after admission, and admission now requires equal epochs on both sides (Host refuses a mismatched hello at
host-kernel.ts:553-577; Client throws on a mismatched accepted handshake atconnection.ts:1580). New Desktop ↔ epoch-198 remote Host gets the terminalRUNTIME_HOST_REMOTE_INCOMPATIBLEerror, which reconnect treats as final (reconnect-lifecycle.ts:200,375), so no storage query is sent and there is no retry loop; old Desktop ↔ new Host is symmetric. This is the usual cost of an epoch bump (the whole remote Host is unavailable from a mismatched Desktop), and it makes the earlier no-backoff re-query concern moot. - Epoch moves 198→199 with the compatible-change declaration removed; the epoch check passes. FYI only: several other open PRs also claim 199, so whichever merges later will need to bump again.
- Since
5af33e59: one commit plus a main merge. The Astryx inventory totals conflict was resolved correctly;task-storage-size.tsxnow observes rows against their nearest scroll container (fixes last round's prefetch nit);storage-footprint.tsis unchanged, so the event-loop blocking fix still holds.
Checks run locally: git diff --check, ASF headers, renderer-architecture, Windows and Astryx inventories (current), locale hygiene, app-shell hooks, merge-tree vs main, full build and Desktop typecheck; storage 28/28, runtime-host 288/288 plus handshake/compatibility 63/63, Desktop 7/7. Not run: a real epoch-198 Host binary, Electron UI, full npm test.
Automated review notice: This comment was posted by an automated review agent. It is not an independent human review and does not replace one.
Astro-Han
left a comment
There was a problem hiding this comment.
Approved at @Astro-Han's explicit request: two independent automated reviews of this head found no blocking issues, the earlier old-Host disconnect is resolved by the epoch 199 handshake gate, and CI is green.
Resolve two conflicts: - packages/runtime-host/src/protocol/index.ts: both sides claimed compatibility epoch 199 (main: storage usage query operations from apache#5832; this branch: UsageQuery.callKinds allowlist). Keep main's 199 entry and re-assign the callKinds entry to 200, bumping RUNTIME_HOST_COMPATIBILITY_EPOCH to 200 and updating the comment's rejected-peer reference from epoch-198 to epoch-199. - apps/desktop/src/main/__tests__/session-inspector-usage-stats.test.ts: semantic union. Keep main's new "carries rounded inspector durations into the next unit" test (apache#5858) in its original position and keep all four callKinds/main-summary tests added by this branch; no assertions removed from either side.
Part of #5825 (M1, storage visibility). Read-only: no data is changed, vacuumed or deleted.
Why
Maka gives users no way to see how much disk it uses or which tasks account for it. Step 1 of the task lifecycle plan (#5776) starts by making that visible. The Host measures and Desktop presents.
What changes
Runtime Host protocol (epoch 198 → 199)
storage.usage.query{}.{ measuredAt, totals, reclaimableBytes, worktreeCount }, with totals of kinddatabase,artifacts,context_offloadandmemory.PRAGMA, or a small metadata sum. Nothing scans the large per-session tables.storage.usage.sessions.querysession_id-leading index, and the query yields to the event loop after every statement.host.resources.query. They return numbers only, never paths.Storage
storage-footprint.tsis a read-only reader on the sharedruntime.sqlitelease. It adds no triggers or write-path counters (tracking(perf): bound storage and protocol read/write amplification #5038).-wal/-shm/-journal) is now a shared helper, and the long-term memory store reuses it. The Subagent worktree directory name is exported from its owner rather than copied.Desktop
formatBytesin@maka/uigains GB/TB and an optional locale. Output below 1 GB is unchanged.Why not a full per-kind breakdown of
runtime.sqliteThe first version summed transcript, runtime and usage bytes across whole tables. Benchmarks showed this is O(table): inline rows put every leaf page on the read path, so a ~375 MB table took about 1 s in a single synchronous call on the Host event loop. The breakdown was dropped rather than moved to counters (write amplification, #5038) or to a worker thread (a new pattern here). Per-task sizes stay, bounded by the 25-ID cap.
Verification
protocol(88), operation dispatcher, host composition and session collaboration authority;protocol-epoch-check(198 → 199),check-renderer-architecture --strict-base, locale hygiene and the Astryx surface inventory.runtime.sqlite, totals take 13 ms and a 25-ID session query takes 90 ms, with the unknown ID omitted.Running app: I checked the running dev app (
dev:worktree, zh-CN) against a disposable copy of a real workspace. The Data page shows the breakdown, and archived tasks show their sizes. The dev log had no errors.Not verified: per-session timing on a multi-GB database.
Open questions (also on #5825)
AI use
Implemented and reviewed with Claude Code. Commits carry a
Co-Authored-Bytrailer.🤖 Generated with Claude Code