Skip to content

fix(driver-sql, driver-turso): reclaimSpace() returns the whole SQLite freelist, not one page per call (#20106) - #20425

Merged
objectstack-fleet[bot] merged 7 commits into
mainfrom
claude/issue-20106-reclaim-space-full-freelist
Sep 28, 2026
Merged

objectstack-fleet[bot] merged 7 commits into
mainfrom
claude/issue-20106-reclaim-space-full-freelist

Conversation

@objectstack-fleet

@objectstack-fleet objectstack-fleet Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #20106
Clause-②: no

reclaimSpace() now returns the whole SQLite freelist on every face. Every reading below comes from a second connection. The file size is read beside it. Each run has an empty-freelist control. Measured head: 0adf65deb.

What was wrong

SQLite's incremental-vacuum program frees one page per sqlite3_step(), and it yields a result row with no columns for each page. A caller that steps it once frees one page. Two clients stepped it once:

  • SqlDriver on better-sqlite3, and TursoDriver in local mode, which is the same code. knex's better-sqlite3 client runs a statement that declares no result columns with Statement.run(), and run() steps once. Statement.reader is false for this pragma.
  • TursoDriver in remote mode. The libSQL client's execute() steps the statement once and leaves it unfinished. The page it freed never reached the file. The unfinished statement also held the connection's implicit transaction open, so a later write on that connection never reached the file either.

What changed

  • packages/drivers/driver-sql/src/sql-driver.ts, SqlDriver.reclaimSpace: when the knex client is better-sqlite3, the pragma runs through that binding's own exec() on the pooled connection. exec() steps every statement until SQLite reports done. Every other SQLite client stays on knex.raw. sql.js steps every PRAGMA to the end in driver-sqlite-wasm's dialect (measured). knex's node-sqlite3 client uses Database.all() (read from knex's source; the binding is not installed here).
  • packages/drivers/driver-turso/src/turso-driver.ts, the remote reclaimSpace route: it reads PRAGMA freelist_count through the raw door first. That read also connects the transport lazily, as every remote door does. When the count is 0 nothing more is sent. Otherwise the vacuum runs through the client's executeMultiple(). Both statements keep the raw door's envelope (DATABASE_ERROR / 500).
  • packages/drivers/driver-turso/README.md: the sentence on the remote reclaimSpace() now says what the route sends.
  • Tests in all three driver packages, and .changeset/20106-reclaim-space-full-freelist.md (patch for driver-sql and driver-turso).

Readings: freelist / page count from a second connection

The fixture fills about 300 pages and deletes them. The control is an empty freelist.

face before main (789b2ae54) after one call this PR after one call
SqlDriver (better-sqlite3, WAL) 300 / 304 299 / 303 0 / 4
SqliteWasmDriver (sql.js image) 300 / 304 0 / 4, already complete 0 / 4
TursoDriver local (better-sqlite3, WAL) 300 / 304 299 / 303 0 / 4
TursoDriver remote (@libsql/client, file:) 300 / 304 300 / 304 (the issuing connection read 299) 0 / 4
control, every face 0 / 4 0 / 4 0 / 4

File size in bytes:

  • Wasm image and remote file: database (rollback journal): 1,245,184 → 16,384 with this PR. On main the remote file stayed at 1,245,184, even after disconnect.
  • The two WAL faces, after disconnect: main 1,241,088, this PR 16,384.

Raw clients, 300 free pages, read from a second connection:

  • better-sqlite3. Statement.run('PRAGMA incremental_vacuum') → 299. incremental_vacuum(600) through run() → 299 too. db.pragma() and db.exec() → 0.
  • @libsql/client file:.
    • execute() → 300. The issuing connection read 299, and the file was unchanged after close().
    • execute('PRAGMA incremental_vacuum(600)') → 300.
    • executeMultiple() → 0.
    • batch([...], 'write') → throws SQLITE_BUSY: cannot commit transaction - SQL statements in progress.
  • sql.js. A prepare-and-step loop (the wasm dialect's branch) → 0. One step → 299.

Remote face, write after reclaimSpace(): a create() made after the call is read back by the issuing connection. On main a second connection counted 0 rows, and after disconnect it still counted 0: the row was lost. With this PR it counts 1, and 1 after disconnect.

Cost at 25,660 free pages through knex on better-sqlite3 (shared box, so read the ratio):

  • exec(): 245 ms.
  • N single-page raw() calls in one transaction: 1,352 ms.
  • N autocommitted raw() calls: 1,606 ms.
  • One raw() call (main): 1.3 ms, for one page.

The dispatch's mechanism hypotheses

  • H1: confirmed on 789b2ae54: if (!this.isSqlite) return; await this.knex.raw('PRAGMA incremental_vacuum');.
  • H2: the first half is confirmed: one execution steps once and frees one page. The second half is falsified. An explicit page count does not help, because incremental_vacuum(600) freed one page through both better-sqlite3 run() and libSQL execute(). The count is a ceiling, not what ends the loop. The statement has to run to completion.
  • H3: partly falsified. SqliteWasmDriver and local TursoDriver do inherit SqlDriver.reclaimSpace. But the defect lived in the client binding, not in the method's text, and the wasm face was already complete on main (its dialect steps every PRAGMA to the end). For this method there are three independent implementations:
    • better-sqlite3 through knex: two faces, both broken, fixed at the SqlDriver seam;
    • the sql.js dialect: one face, already correct, now pinned;
    • the libSQL remote route: one face, broken, fixed in turso-driver.ts.
  • H4: confirmed, and worse than a reading. Through libSQL execute(), a second connection saw nothing land: not the freelist, not the file size, and not a later write on the same connection.

Does os db clean reach this method?

No. packages/cli/src/commands/db/clean.ts never calls reclaimSpace(). It runs driver.execute('PRAGMA auto_vacuum = INCREMENTAL'), then driver.execute('VACUUM'), then disconnects. A full VACUUM finishes in one step. Measured through SqlDriver.execute on a legacy file (auto_vacuum 0) with 300 free pages:

  • freelist 300 → 0;
  • auto_vacuum 0 → 2;
  • file 1,236,992 → 12,288 bytes.

It is not edited here.

The remote face: measured and not measured

Measured over a libSQL file: client only (@libsql/client 0.17.4, libsql 0.5.29), plus a scripted client for the call order and the refusal envelope.

Not measured, because there is no live server: what a hosted libSQL / Turso server does. That covers whether it accepts PRAGMA freelist_count and PRAGMA incremental_vacuum, and how it steps execute() against executeMultiple() (a hrana sequence request). The route adds one read round trip when there are free pages. When the freelist is empty it sends no write.

Tests

  • driver-sql/src/sql-driver-sqlite-reclaim-space.test.ts (new), 4 cases:
    • WAL: freelist 0 and page count equal to before minus free, read from a second connection; file size equal to pages times page size after close;
    • DELETE journal: the same, and the file shrinks while the driver is open;
    • the empty-freelist control;
    • the pooled connection is handed back.
  • driver-sqlite-wasm/src/sqlite-wasm-reclaim-space.test.ts (new), 2 cases: the persisted image, read by a fresh sql.js database, plus the control.
  • driver-turso/src/turso-remote-inherited-members.test.ts: the old pin read the freelist from the issuing client and asserted toBeLessThan(before), which a one-page, connection-local drop passed. It now reads a second client and asserts equality. Added:
    • the control on both faces;
    • the remote file shrinking while open;
    • a write after the call landing;
    • the call order (count first, nothing more on 0, vacuum through executeMultiple);
    • the DATABASE_ERROR / 500 envelope when either statement is refused.

Full suites at 0adf65deb:

  • driver-sql: 195 files passed, 11 skipped; 3,224 tests passed, 178 skipped.
  • driver-turso: 73 files; 1,927 tests passed, 16 skipped.
  • driver-sqlite-wasm: 36 files; 658 tests passed.
  • typecheck is green for all three. Each package's tsconfig includes its tests; the typecheck read the turso test file and caught an Array.prototype.at before it landed.

Ablations (committed state first; every leg went through scripts/ablation-replace.mjs, and the restore was proven blob-equal to HEAD):

  • A: better-sqlite3 routed back to knex.raw. The driver-sql test goes red on WAL and DELETE ({ freelist: 299, pages: 303 } against { freelist: 0, pages: 4 }); the control and the pool case stay green. With driver-sql rebuilt and the marker proven in dist/ by ablation-dist-preflight, the turso local face goes red ({ freelist: 39, pages: 47 }) and the wasm test stays green (2/2). That is the H3 reading. The restore was rebuilt, and --absent proved dist/ and the tree clean.
  • B: the knex.raw arm made a no-op, dist/ rebuilt, marker proven. The wasm test goes red ({ freelist: 300, pages: 304 }) and its control stays green. The restore was rebuilt and --absent passed.
  • C: the remote route sent back through execute(). 5 turso cases go red: second-connection freelist 40 against 0, the file size case, the write after the call (0 against 1), the call order, and the vacuum refusal. The other 77 stay green.

Gates

node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack at 0adf65deb derives 63 commands. All 63 were run, and every exit code was recorded before any pipe: all 63 exited 0.

On the first run at 155234b97:

  • check:query-options-erasure was red: the test surface grew 236 → 237 from an as any on a query bag in the new tests. The bags are now typed, and it holds at 236.
  • check:dual-build-cjs-loads, check:lean-entry-closure and check:type-check-debt refused with PREREQUISITE NOT MET (exit 3). They passed after turbo run build over ./packages/* and ./packages/*/*.

--ran: 63 derived, 63 run, 0 NOT-MEASURED, 0 UNRUN.

check:driver-conformance: before, at 789b2ae54, 50 covered cells and 0 in the DEBT ledger (0 exempt). After, at 0adf65deb, the same: 50 covered, 0 DEBT.

pnpm lint is CI's run. The narrowed run here: eslint --no-inline-config --format json over the 5 changed TypeScript files reports 5 files, 0 errors and 0 warnings, with no file reported as ignored. eslint.config.mjs sets no parserOptions.project and registers no typed rule, so a diff cannot move the verdict of an untouched file.

Acceptance notes


Generated by Claude Code

…r-sqlite3

reclaimSpace() stepped the statement once through knex's Statement.run(),
which frees one freelist page per call. Drive the better-sqlite3 binding
through its own exec(), which steps until SQLite reports done.

Claude-Session: https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN
Co-authored-by: Claude <noreply@anthropic.com>
The remote route sent PRAGMA incremental_vacuum through the libSQL client's
execute(), which steps it once and leaves it unfinished. Read the freelist
count through the raw door, then run the vacuum with executeMultiple().

Claude-Session: https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN
Co-authored-by: Claude <noreply@anthropic.com>
…SQLite face

Read the freelist, page count and file size from a second connection on
driver-sql, driver-sqlite-wasm and both driver-turso faces, with an empty
freelist control; pin that a write after the remote call lands. Update the
remote route's README sentence and add the changeset.

Claude-Session: https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN
Co-authored-by: Claude <noreply@anthropic.com>
Array.prototype.at is outside this package's ES2020 lib.

Claude-Session: https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN
Co-authored-by: Claude <noreply@anthropic.com>
Drops the as-any erasures check:query-options-erasure counts.

Claude-Session: https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN
Co-authored-by: Claude <noreply@anthropic.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation tests tooling labels Sep 28, 2026
@github-actions

Copy link
Copy Markdown
Contributor

📓 Docs Drift Check

This PR changes 2 package(s): @objectstack/driver-sql, @objectstack/driver-turso, touching 3 documentable anchor(s). ⚠️ 1 changed file(s) yielded no anchor (packages/drivers/driver-turso/README.md), so the pages documenting them are NOT COVERED by this run — this is not a clean bill of health for those files.

8 hand-written doc(s) NAME something this change touched and may need an implementation-accuracy re-verification:

  • content/docs/data-modeling/drivers.mdx (via SqlDriver (symbol, a top-level class), TursoDriver (symbol, a top-level class), reclaimSpace (symbol, a method of class SqlDriver; a method of class TursoDriver))
  • content/docs/data-modeling/index.mdx (via SqlDriver (symbol, a top-level class))
  • content/docs/permissions/tenant-audit-census.mdx (via SqlDriver (symbol, a top-level class))
  • content/docs/plugins/packages.mdx (via SqlDriver (symbol, a top-level class), TursoDriver (symbol, a top-level class))
  • content/docs/protocol/kernel/index.mdx (via SqlDriver (symbol, a top-level class))
  • content/docs/protocol/kernel/lifecycle.mdx (via SqlDriver (symbol, a top-level class))
  • content/docs/protocol/objectql/query-syntax.mdx (via SqlDriver (symbol, a top-level class))
  • content/docs/protocol/objectql/types.mdx (via SqlDriver (symbol, a top-level class))

⛔ 2 release-owned page(s) also name something this change touched. These are read-only:

  • content/docs/releases/v17/17-0.mdx (via SqlDriver (symbol, a top-level class))
  • content/docs/releases/v17/17-5.mdx (via SqlDriver (symbol, a top-level class), TursoDriver (symbol, a top-level class))

content/docs/releases/ is RELEASE-OWNED (AGENTS.md "Documentation Guardrails"): release
notes are written centrally at release time, and a code PR that edits them is the exact PR
that guardrail exists to stop. They are still audited — read-only. If one of them is actually
wrong, file an issue or open a dedicated docs-only PR; do not edit it here.

What this run could not see
  • 1 changed file(s) yielded no anchor (packages/drivers/driver-turso/README.md) — pages documenting those are invisible to this run
  • the SDK route bridge reached 54 of 206 client-bound route-ledger rows — the other 152 have no registrar path: tail to select them, so pages documenting THEIR client methods cannot appear above, on this or any run. Of those 152: 0 are remediable by widening that discovery convention (an in-repo file declares the path; the convention did not scan it); 55 are structural — on a ledger where NOT ONE row is declared in-repo, so no discovery change reaches them at any price; 97 are undecided (no in-repo declaration, on a ledger that has other in-repo registrars — absence and an unreadable spelling are not distinguishable here). The rows themselves: node scripts/docs-audit/affected-docs.mjs --bridge-coverage
  • a page that states a rule by its inputs shares no identifier with the emitter that implements the rule, so an emitter-only diff cannot list it — not on this run and not on any run. Measured on fix(driver-sql): emit varchar(maxLength) for a text field a declared index keys on #11430: content/docs/protocol/objectql/types.mdx documents the text-family column mapping by the ObjectQL type names it maps FROM (text / textarea / html) while the diff changed createColumn; it went unlisted, and it was the page that diff falsified, in four places. No shared token exists to detect this on, so a rule your change carries has to be re-read by hand in the pages that restate it.
  • a key NAME is not a key, so the hand re-read the line above prescribes can land on the wrong schema. The same spelling is authorable on one governed type and a [REMOVED] tombstone on another for each of active, aria, joins, objects, template, tools and version (censused on [finding] tools is a key on BOTH AgentSchema (tombstoned, dead) and SkillSchema (live, cloud-attested), so a name-based search attributes skill examples to the agent key — it produced a false stop-the-line alarm on PR #19059 #19093 over the liveness ledger's governed types, top-level keys); nothing in a search result distinguishes the two, so a grep hit on a LIVE example reads as evidence about the DEAD key. Measured on fix(spec): the agent.tools liveness row says dead — it claimed live on a key the schema tombstoned #19059: content/docs/ai/agents.mdx was reported as contradicting the agent.tools tombstone over its tools: example at :161, which is inside the defineSkill({ block opened at :155 — the page was already correct. Settle ownership by PARSING the value against both schemas, never by the name: that literal PASSES SkillSchema, and as an AgentSchema it FAILS at tools with the tombstone prescription. ⛔ These names are not the whole class — a key retired through a .strict() guidance map leaves no tombstone in the walked shape and none of them here (tool.category, live as AIToolDefinition.category).

Coarse fallback — 13 page(s) merely mention a changed package (the pre-#9192 predicate, kept for the deliberately-wide backstop): node scripts/docs-audit/affected-docs.mjs --json 0283cb924a5e29f22fdf0db60bcaedeab3ead3f3 → packageMentionDocs.

Which tree this was computed on

This run read content/docs from 39fcf4ed91bdd2a8621ced2f5250e39f26887a7c — the merge of head 0adf65debb127f52660624df7ac9af06b752c1e0 into base 0283cb924a5e29f22fdf0db60bcaedeab3ead3f3, which is what actions/checkout gives a pull_request run. Not the PR head.

A worktree cut from an older main holds a different content/docs, so re-deriving there can legitimately return a different list — that is a different tree, not a wrong row. To answer on the same tree:

# while this PR is open — GitHub drops the merge commit once it closes
git fetch origin 39fcf4ed91bdd2a8621ced2f5250e39f26887a7c && git checkout 39fcf4ed91bdd2a8621ced2f5250e39f26887a7c
# afterwards, rebuild it from the two parents, which stay fetchable
git fetch origin 0283cb924a5e29f22fdf0db60bcaedeab3ead3f3 0adf65debb127f52660624df7ac9af06b752c1e0 && git checkout -B drift-repro 0283cb924a5e29f22fdf0db60bcaedeab3ead3f3 && git merge --no-ff 0adf65debb127f52660624df7ac9af06b752c1e0

node scripts/docs-audit/affected-docs.mjs --json 0283cb924a5e29f22fdf0db60bcaedeab3ead3f3

⚠️ That checkout carried uncommitted changes, so the commit above does not fully identify what was read.

Advisory only, and a precision-first one (#9192): a page is listed because it names a
symbol, wire route or SDK method this diff touched — not because it mentions a changed
package. Each row says which anchor put it there, so a wrong row is reportable rather than
merely annoying. To re-verify, run the docs-accuracy-audit workflow scoped to these files:
node scripts/docs-audit/affected-docs.mjs 0283cb924a5e29f22fdf0db60bcaedeab3ead3f3 → pass the list as
args.docs, on the commit named under Which tree this was computed on.

@objectstack-fleet

Copy link
Copy Markdown
Contributor Author

Contract review

Served-tier: CONTRACT_REVIEW_TIER
Head-sha: 0adf65debb127f52660624df7ac9af06b752c1e0
Local-runs: none

① Derived judgments

Inputs read: card #20106 (body and all 5 comments, 5827314394 through 5868473417), PR #20425 (body, 7-file list) and git diff origin/main...origin/claude/issue-20106-reclaim-space-full-freelist at 0adf65deb (merge base 50e273fd7; +410 / −35 over 7 files, equal to the API file list), plus the vendored bindings the diff's own comments cite (knex 3.3.0, @libsql/client 0.17.4, the wasm dialect on main). Contract: IDataDriver.reclaimSpace (packages/spec/src/contracts/data-driver.ts 451–457: reclaim free space after bulk deletions, ADR-0057 §3.4; optional, best-effort after every sweep that deleted rows) and ADR-0057 §3.4 (the Reaper issues PRAGMA incremental_vacuum after a sweep). Neither is edited.

  1. SqlDriver.reclaimSpace, the better-sqlite3 arm — RIGHT. knex's Client_BetterSQLite3._query (vendored, lines 51–57) runs a statement whose reader is false through Statement.run(): one step, one freed page. Database.exec() steps to done. The arm keys on client.driverName === 'better-sqlite3' (the dialect's declared driverName, line 98), takes client.acquireConnection() and hands it back in finally. The knex.raw arm stays for every other SQLite spelling isSqlite admits (sqlite3, sqlite, better-sqlite3): knex's node-sqlite3 client dispatches a raw statement to all (vendored line 149; that binding is not a dependency of driver-sql, so the PR rightly calls it read, not measured), and the wasm dialect (driverName: 'wasm-sqlite') runs every PRAGMA through while (stmt.step()) — complete on main, so the wasm face needs and gets no source change; SqliteWasmDriver overrides isSqlite to true, so the method reaches that arm. Faces delivered: SqlDriver on better-sqlite3; TursoDriver local, whose knex config is client: 'better-sqlite3' (turso-driver.ts 1569 / 1576 / 1584) and so takes the same arm; SqliteWasmDriver, pinned unchanged.

  2. TursoDriver.reclaimSpace, the remote route — RIGHT. Order: transport.execute('PRAGMA freelist_count') (RemoteTransport.execute calls ensureConnected(), the lazy connect every remote door uses, and returns result.rows; a libSQL Row is index-addressable, so rows[0]?.[0] reads the scalar) → return on 0 → transport.getClient()!.executeMultiple('PRAGMA incremental_vacuum'), non-null once ensureConnected resolved. Both statements sit in rawStatementFault → DATABASE_ERROR / 500, pinned per refusal with a scripted client. Mechanism verified in the vendored @libsql/client file: implementation: executeStmt calls sqlStmt.raw(true), which throws for a column-less statement, so it falls to sqlStmt.run(args) (one step); executeMultiple is db.exec(sql). assertRemoteTransactionUnsupported and the local fall-through to super are unchanged. Edge, noted not flagged: a mis-shaped or empty count answer yields NaN, which is not === 0, so the vacuum still runs — the safe direction. Option A's named cost stands: a server refusing only the read now refuses the route at the read; LifecycleService.sweep() wraps reclaimSpace() in try/catch and logs a warning (lifecycle-service.ts 710–717), the same degradation a vacuum refusal had on main.

  3. The remote face over a hosted server — NOT MEASURED, and nowhere stated as measured. The PR body's readings table labels its remote row @libsql/client, file:; the body's "measured and not measured" section, the changeset and the route's docstring each say what a hosted libSQL server does with either call is not measured. The lead sentence "on every face" is a design claim that the same body bounds two sections later; no reading is attributed to a hosted server.

  4. The README bullet (packages/drivers/driver-turso/README.md, the one reclaimSpace() bullet; 5 lines out, 7 in, nothing else in the file) — TRUE against the code: the route reads PRAGMA freelist_count; when there are free pages it runs PRAGMA incremental_vacuum through executeMultiple(); free pages come only under auto_vacuum=INCREMENTAL (local connect sets it, sql-driver.ts 6857; the remote face does not — the unchanged half of the bullet); either refusal answers DATABASE_ERROR / 500. "To completion there" is what executeMultiple exists for and what the file: reading shows; the sentence claims no reading of a hosted server, so it states nothing as measured that was not.

  5. The accept-set the tests pin — RIGHT, and the tightening the card asked for. Every reading comes from a SECOND connection (a read-only better-sqlite3 knex; a fresh sql.js database over the persisted image; a fresh libSQL client), with { freelist: 0, pages: before minus free } equality replacing the old issuing-client toBeLessThan; an empty-freelist control on every face; file size checked (DELETE journal while open, WAL after close, remote file: while open); the pooled connection handed back; the remote post-call write reaching a second connection; the call order and both refusal envelopes. No production behaviour is loosened anywhere.

  6. os db clean (triage note 3) — answered correctly and not edited: packages/cli/src/commands/db/clean.ts 101–102 runs PRAGMA auto_vacuum = INCREMENTAL then VACUUM through driver.execute; packages/cli contains no call to reclaimSpace.

  7. Public surface: both reclaimSpace signatures unchanged; the spec contract untouched; the remote face sends one extra read when pages are free and nothing when the freelist is empty; a supplied client must carry executeMultiple, which the @libsql/client Client type already declares. No accept-set is widened; "reclaimed" in the lifecycle report stays best-effort, as the contract says.

② Semver level

.changeset/20106-reclaim-space-full-freelist.md: @objectstack/driver-sql: patch, @objectstack/driver-turso: patch. Matches what the diff publishes: the two packages with source changes ship a behaviour fix under an unchanged signature and an unchanged contract; driver-sqlite-wasm changes only a test file and publishes nothing (Check Changeset: success). Clause-②: no — right: the change brings the method to what IDataDriver.reclaimSpace and ADR-0057 §3.4 already declare, and nothing in packages/spec or the ADR moves.

③ Boundary flags

Check-runs on 0adf65deb, final read 2026-09-28T11:05:20Z: 40 runs, 39 completed — 34 success, 5 skipped (Auto Label and Check PR Size on the 10:57 re-run, Console Pin Gate, Build Docs, Packed-tarball smoke opt-in), 0 failure. Green: Build Core; Test Core 1/6, 2/6, 3/6, 4/6, 5/6; Type Check source gates, workspace, consumer gates, debt ledger and TypeScript Type Check; Lint & Repo Gates; Temporal Conformance (live PG + MySQL); Dogfood Verify CLI and Dogfood Regression Gate 1/3, 2/3, 3/3; Check Changeset; Governed Surface Queue Guard; Flag docs affected by code changes; Check Documentation Links; the card/branch/single-writer/part-of guards (both runs). Test Core (6/6), the last run, completed success at 2026-09-28T11:07:33Z (read from the poll armed on the same endpoint): 40 runs, 40 completed, 35 success, 5 skipped, 0 failure. Every gate family on the head is green.

Implemented-by: claude/issue-20106-reclaim-space-full-freelist
Reviewed-by: session_01N8TPEsoJxPsdSdNKGnNGEN

VERDICT: PASS

@objectstack-fleet
objectstack-fleet Bot marked this pull request as ready for review September 28, 2026 11:11
@objectstack-fleet
objectstack-fleet Bot added this pull request to the merge queue Sep 28, 2026
Merged via the queue into main with commit e01d347 Sep 28, 2026
43 checks passed
@objectstack-fleet
objectstack-fleet Bot deleted the claude/issue-20106-reclaim-space-full-freelist branch September 28, 2026 11:35
veigajoao pushed a commit to veigajoao/objectstack that referenced this pull request Sep 29, 2026
…mote mode, and sync with no syncUrl (objectstack-ai#20200) (objectstack-ai#20447)

Fixes objectstack-ai#20200
Clause-②: no (narrowing)

The `Clause-②` line above is the amended claim's (comment 5869128452,
unchanged in 5871001040), copied as it stands. The changeset carries the
same value.

Session `session_01N8TPEsoJxPsdSdNKGnNGEN` (PM dispatch, `domain:engine`
seat 1, mode:subagent), branch
`claude/issue-20200-turso-remote-sync-refused`. The branch starts at
`e01d34730`, which already contains PR objectstack-ai#20422 and PR objectstack-ai#20425. It was then
merged with `origin/main` at `87c37aec1` in a true merge commit. **Every
reading below was taken at head `2242ad513`** unless it says otherwise.

## What changes

`new TursoDriver(...)` refuses two configurations it used to build while
ignoring one of their keys. Both are refused with `VALIDATION_ERROR` /
400, before `super()`, after the constructor's three existing refusals,
so no configuration they already refuse gets a different message:

- **`syncUrl` under a forced `mode: 'remote'`.** A remote url beside
`syncUrl` with no `mode` was already refused, as a replica on a remote
url (`localEngineDefect`), so only a forced remote mode reaches this
refusal.
- **`sync` with no `syncUrl`**, in every mode: local, replica and
remote. An empty `syncUrl` counts as unset, which matches `detectMode`,
`connect()`, `sync()` and the spec's refinement.

Each refusal's message is `@objectstack/spec`'s `TursoConfigSchema`
issue message for that key, byte for byte. Per the seat's ruling (option
A of the report `5869079328`), each is a module-level constant in
`turso-driver.ts`: `REMOTE_MODE_SYNC_URL_REFUSAL` and
`SYNC_WITHOUT_SYNC_URL_REFUSAL`. The parity test holds each constant
equal to the schema's issue, read from the built spec dist. No new
`packages/spec` export.

**The message must be true where it is thrown** (seat ruling 2, a
declared cross-lane text edit). The spec's `syncUrl`-under-remote text
said "the turso driver never hands `syncUrl` to the remote client and
runs no sync, so the setting changes nothing". That clause now reads
"the turso driver refuses this configuration when it starts", the form
its sibling refusals in `tursoTransportIssues` use. The rest of the
message is byte-identical. The three copies now read the same: the spec,
the driver's `TursoConfigSchema` mirror and the constructor constant.
The `sync` message is reused unchanged. The ADR-0087 entry
`18.turso-config-transport-mismatch-refused` has its `reason` sentence
on stored rows corrected. It now says the constructor also refuses
`syncUrl` under a forced remote mode and `sync` with no `syncUrl` when
the datasource boots, citing objectstack-ai#20200. Patch round 1 also put its header
comment and its "One more it builds and then ignores" clause in the past
tense, bounded by objectstack-ai#20200. Nothing else in the entry moved.
`check:generated` then proved `src/migrations/registry.ts` stale,
because the registry embeds the entry text, and `check:generated --fix`
regenerated it: a 4-line text diff. No gate refused editing a registered
entry.

## H1: before and after (dist probe, 9 configs)

The same scratch script ran against
`packages/drivers/driver-turso/dist/index.mjs`. Before is at `dbddf02c1`
(origin/main when the first round measured). After is this branch's
build.

| config | before (`dbddf02c1`) | after |
| --- | --- | --- |
| `libsql://` + `mode: 'remote'` + `syncUrl` + `sync` | constructs,
connects, `isSyncEnabled()` true, no interval, `sync()` rejects
`SYNC_NOT_SUPPORTED` | refused, VALIDATION_ERROR / 400, `syncUrl`
message |
| `libsql://` + `mode: 'remote'` + `syncUrl` | the same | refused,
`syncUrl` message |
| `file:` + `mode: 'remote'` + `syncUrl` + `sync` | constructs,
connects, `isSyncEnabled()` true, no interval, `sync()` rejects
`SyncNotSupported("File")` | refused, `syncUrl` message |
| `libsql://` + `mode: 'remote'` + `sync`, no `syncUrl` | constructs,
`isSyncEnabled()` false, `sync()` a no-op | refused, `sync` message |
| `libsql://` (no mode) + `sync`, no `syncUrl` | the same | refused,
`sync` message |
| CONTROL `libsql://` + `mode: 'remote'` | constructs, connects,
`isSyncEnabled()` false | unchanged |
| RIDER `file:` + `mode: 'replica'`, no `syncUrl` | constructs as
`replica`, `isSyncEnabled()` false, `sync()` a no-op | **unchanged (see
the rider section below)** |
| RIDER + `sync`, no `syncUrl` | the same | refused, `sync` message (the
`sync` key, not the rider) |
| `file:` + `sync`, no `syncUrl`, no mode | local, `sync` ignored |
refused, `sync` message |

## H4: what a stored row answers at boot now

`buildTursoDriverConfig`
(`packages/services/service-datasource/src/turso-driver-config.ts`, read
and not edited) forwards a stored row's `syncUrl`, `sync` and `mode`
unparsed. PR objectstack-ai#20199's ADR-0087 entry is semantic, so no D2 conversion
rewrites such a row either. A row stored before PR objectstack-ai#20199 with a remote
url, `mode: 'remote'` and `syncUrl`, or with `sync` and no `syncUrl`,
therefore reaches the constructor as written:

- `factory.create` throws the refusal
(`default-datasource-driver-factory.ts` for open-core,
`turso-driver-factory.ts` for the host default).
- `DatasourceConnectionService` catches it and records the datasource as
`failed-degraded`, with "datasource NAME: connect failed — MESSAGE".
- Under ADR-0062 D5, boot fails fast when objects bind to that
datasource or are routed to it, or when it is boot-critical, unless
`OS_ALLOW_DRIVER_CONNECT_FAILURE` is set. Otherwise it is left
unconnected with a warning.
- A test connection answers `ok: false`, with "Failed to build driver:
MESSAGE".

Before this change the same row booted, reported sync as enabled, and
never synced. No in-repo caller reads `isSyncEnabled()` or calls the
driver's sync outside `@objectstack/driver-turso`'s own tests. The
changeset's FROM → TO paragraph carries these readings and the way out:
drop `syncUrl` / `sync` from a remote config, or use a `file:` url with
the remote in `syncUrl`.

## The rider stays: `mode: 'replica'` on a `file:` url with no `syncUrl`

This configuration still constructs and runs as a plain local database.
The table's H1 rider row shows it: a declared replica that never syncs.
It is deliberately not refused here. A constructor-only refusal would
make construction refuse a configuration both `TursoConfigSchema` copies
accept, which reopens objectstack-ai#19977's defect class in reverse (a datasource
that authors clean and fails at boot). The parity table would go red on
its row "file: under a forced mode: 'replica'", and
`turso-driver-unrecognised-url-refusal.test.ts` pins `FILE: + mode
'replica', no syncUrl` as accepted. Closing it needs the spec half and
the constructor half together, and the spec's accept set is outside this
card, so the seat files it as its own card. The new test file pins the
rider as still accepted, so whoever closes it moves that pin on purpose.

## Tests

| suite at `2242ad513` | result |
| --- | --- |
| `@objectstack/driver-turso` vitest, whole package | 75 files · 2014
passed · 18 skipped · exit 0 |
| `@objectstack/driver-turso` typecheck (`tsc --noEmit`) | exit 0;
`--listFilesOnly` on the pre-merge commit shows both touched test files
in the program |
| `@objectstack/spec` vitest `--project local`, 3 shards | 564 files ·
16640 passed · 1 todo (6026 + 5141 + 5473), exit 0 on each shard |
| `@objectstack/spec` typecheck (tsc + scripts + `check:test-typecheck`)
| exit 0 |

The 18 skips are the parity table's forced-mode rows for the mirror,
which strips `mode`: 16 before, plus the two new forced-mode `sync`
rows.

- **`spec/turso-config-constructor-parity.test.ts`:** the four `inert`
rows flip to `ctor: 'refuse'` (three `syncUrl`, one `sync`). Four `sync`
rows are added: a remote url, a forced remote, a forced replica on
`file:`, and an empty `syncUrl`. The `inert` floor moves from at least 4
to exactly 0, with floors added for `syncUrl` (at least 3), `sync` (at
least 5) and the sync-key refusals (at least 8). A new table,
`SYNC_KEY_REFUSALS`, asserts for each of those 8 rows that the
constructor's `error.message` equals the spec issue's message.
- **`turso-driver-ignored-sync-key-refusal.test.ts` (new):** both
refusals are asserted as the envelope (`code` + `status`) plus the
message's first sentence, across `libsql://`, `https://`, `file:`,
`:memory:`, forced remote, forced replica and an empty `syncUrl`.
`createTursoDriver()` is covered too. Controls: forced remote with no
`syncUrl` connects and reports sync off; a `file:` replica beside
`syncUrl`; `sync` beside `syncUrl` under forced `mode: 'local'`; the
rider.
- **`packages/spec/src/data/driver/turso.test.ts`:** one assertion and
its test title pinned the old clause ("never hands `syncUrl` to the
remote client and runs no sync"). They now pin the new clause. This file
is not named in the claim's surface; the change is the mechanical
consequence of ruling 2's text edit (see Deviations).

**Reverse verification**, via `scripts/ablation-replace.mjs` from the
committed state (turso-driver.ts blob `afe3ad31`). The driver tests
import `../turso-driver` from source, so no build is involved.
Directions were predicted before each run, and all three matched:

1. `if (mode === 'remote' && config.syncUrl) {` became `if (false && …)
{`, mutation landed (anchor 1 → 0, blob `afe3ad31` → `5f9aca88`). **11
failed** / 169 passed: exactly the 5 `syncUrl` cases in the new file,
plus the 3 `syncUrl` rows' constructor verdicts and their 3 message
pins. Restored: blob == HEAD, `git diff HEAD` empty.
2. `if (config.sync && !config.syncUrl) {` got the same mutation (blob →
`32b6a28c`). **16 failed** / 164 passed: the 6 `sync` cases, plus 5
constructor verdicts and 5 message pins. Restored the same way.
3. One byte of the copy: a doubled space inside the `syncUrl` constant,
after its first sentence (blob → `084c8d96`). **3 failed**: exactly the
3 `syncUrl` message pins, while every verdict and first-sentence case
stayed green. This proves the byte-equality pin is what holds the copy.
Restored the same way.

## Gates

`node scripts/pm/dispatch-gates.mjs --commands --repo
objectstack-ai/objectstack`, run after the last commit at `2242ad513`,
lists 9 paths vs merge base `87c37aec1` and **91 commands**. Every
command ran, with its exit code written to disk before any pipe. `--ran`
reconciliation reads `91 derived famil(ies) accounted for — 89 run, 2
NOT-MEASURED (2 DERIVED from a recorded exit 3)`, with 0 UNRUN.

- **NOT MEASURED (exit 3, PREREQUISITE NOT MET):**
`check:dual-build-cjs-loads` (dozens of workspace packages have no
`dist/`) and `check:type-check-debt` (it needs a whole-workspace build).
Both need a full-workspace build, which is CI's run.
- **Run twice:** `check:doc-formula-expressions` and
`check:lean-entry-closure` first answered exit 3. They are exit 0 after
building `@objectstack/lint` and `@objectstack/objectql` with their
closures.
- **Notable readings:**
- `check-adr-0087-registration` → `not-required (already-registered)`
for `turso-config-transport-mismatch-refused`, clause-② narrowing,
BREAKING, bang;
  - `check-changeset-no-major` exit 0;
  - `check:migration-registry` → `registry.ts is current`;
- `check:driver-conformance`, `check:nul-bytes`, `check:doc-authoring`,
`check:test-source-alias`, `check:cross-package-test-inputs`,
`check:api-surface`, `check:authorable-surface`, `check:docs`,
`check:spec-changes` and `check:upgrade-guide` → exit 0.
- **Roster families under a touched directory, also run:** five
artifact-roster families whose roster sits under a directory this diff
touches, all exit 0: `check-changeset-fixed`, spec
`check:meta-url-spelling`, `check:authz-resolver`,
`check:error-code-casing` and `check:filter-alias-parity`.
- **Narrowed lint:** `eslint --no-inline-config --format json` over the
8 changed TS files reports 8 files, 0 errors and 0 warnings. The
population is `eslint.config.mjs`'s lint object, `files:
['**/*.{ts,tsx,mts,cts,js,jsx,mjs,cjs}']`; the changeset `.md` is
outside it. Invariance holds because that config never enables
type-aware linting (no `parserOptions.project`, per its own header), so
this diff cannot move any untouched file's verdict. The full `pnpm lint`
is CI's.
- **Not run locally, left to CI:**
  - the whole-workspace type-check lanes;
  - the Test Core, Dogfood and Build Core jobs;
- downstream suites of `@objectstack/service-datasource`,
`@objectstack/runtime` and `@objectstack/cli`. Their turso fixtures go
through `buildTursoDriverConfig` or a capturing constructor and never
build the real driver. The downstream real-driver constructions are both
in dogfood's `date-bucket-parity-turso.test.ts`: `new TursoDriver({ url:
':memory:' })`, and at `:144` a remote `{ url:
'libsql://test-db.turso.io', authToken }`. Neither carries a sync key,
so neither refusal touches them.

**Driver-conformance ledger (`lanes/engine.md`):**
`check:driver-conformance` read `OK — 50 covered cell(s), 0 in the DEBT
ledger, 0 exempt` both before (`dbddf02c1`) and after (`2242ad513`).
`driver-turso` is `ok` on all 10 case-sets both times. No movement.

## Deviations (declared)

- `packages/spec/src/data/driver/turso.test.ts`: one assertion and one
test title, rewritten because ruling 2's text edit makes the old pin
false. It is outside the claim's listed surface ("nothing else in
`packages/spec`"). Without it the spec suite goes red.
- `turso-driver.ts` outside the constructor body. It carries the two
module constants and a small throw helper (`refuseIgnoredSyncKey`)
beside `refuseSuppliedClientTimeout`, as ruling 1 prescribes. It also
extends the TSDoc on `TursoDriverConfig.syncUrl` and `.sync` so the
declared config names the two new refusals. No other region of the file
changed.

## Acceptance notes

- **The mirror's dormant copy is now kept equal by this diff.** The
driver mirror (`src/spec/turso.zod.ts`) strips `mode`, so its copy of
the `syncUrl`-under-remote text is still unreachable through a parse,
and no test can hold it equal. This diff applies the same one-clause
edit there, so the three copies read identically today.
- **The spec wording item from the report is done here** (ruling 2).
- **Comments that described the ignored sync keys: corrected in patch
round 1** (the amended claim 5871001040). Every comment in this diff's
files that described `syncUrl` under a forced remote mode, or `sync`
with no `syncUrl`, as constructed and ignored now says what holds since
objectstack-ai#20200 (list below). A `git grep -n -i ignore` over the touched files
finds no such wording left. What remains is only
`packages/drivers/driver-turso/README.md`'s list of constructor
refusals. It is scoped to the local-engine refusals and does not name
the two new ones: incomplete, not false (the review's reading), and a
docs follow-up. It is not runtime text.
- Loader fixtures in `@objectstack/service-datasource`
(`turso-driver-config.test.ts`, the `mode: 'remote'` + `syncUrl`
key-list case) spell a configuration the real constructor now refuses.
They exercise `buildTursoDriverConfig` only and never construct the
driver, and the package is fenced, so they are unchanged.

## Patch round 1 (text only, head `7052b9111`)

The at-tier review 5870987840 PASSed `2242ad513` and escalated comments
this diff made false. With the `inert` floor at 0, the driver no longer
constructs and ignores anything the schemas refuse. Per the amended
claim 5871001040, those comments are corrected here as text only. No
logic and no test assertion changed.

- **Spec `packages/spec/src/data/driver/turso.zod.ts`:**
  - section heading 1b "refuses or ignores" → "refuses";
- the intro "or constructs and then ignores" → "or, until objectstack-ai#20200,
constructed and then ignored";
- the `syncUrl`-under-remote bullet ("which the driver accepts and then
IGNORES … That arm has no constructor refusal behind it") is now in the
past tense and says the constructor refuses it too since objectstack-ai#20200;
- "Nothing the constructor accepts is refused here, the
`syncUrl`-under-`mode: 'remote'` arm aside" now says that exception has
been a constructor refusal too since objectstack-ai#20200;
- the superRefine comment "What the driver refuses at construction, or
constructs and ignores" → "What the driver refuses at construction".
- **Mirror `packages/drivers/driver-turso/src/spec/turso.zod.ts`:**
section heading 2a "refuses or ignores" → "refuses"; the intro "or
constructs and ignores" → "or, until objectstack-ai#20200, constructed and ignored";
the superRefine comment corrected as in the spec.
- **Spec `turso.test.ts`:** the block comment "(which it constructs and
ignores)" → "(since objectstack-ai#20200 that includes `syncUrl` under a forced `mode:
'remote'`, which it used to construct and ignore)".
- **Parity test:** two comments. The header's "(or constructs and
ignores)" → "(or, until objectstack-ai#20200, constructed and ignored)", and "nothing
it accepts but the declared-and-ignored keys" → "… but an `inert` row's
declared-and-ignored key (none since objectstack-ai#20200)".
- **D3 entry `18.turso-config-transport-mismatch-refused`:**
- its header comment now reads "the one combination the driver used to
build and then ignore (syncUrl under a forced remote mode), which the
constructor refuses too since objectstack-ai#20200";
- its `reason` clause "One more it builds and then ignores … Nothing the
constructor accepts is refused, that key aside" is in the past tense and
bounded by objectstack-ai#20200 ("One more it built and then ignored until objectstack-ai#20200 …";
"Nothing the constructor accepts is refused (at objectstack-ai#19977 that key was the
one exception; since objectstack-ai#20200 there is none)");
- nothing else in the entry moved, and `registry.ts` was regenerated by
`gen:migration-registry`, its hunk equal to the entry's.
- **Left as they are, each true:** "would ignore" (a conditional, naming
what the refusals prevent), the `inert` mechanism's own definition (with
none today), the "before" measurement table in the new test file, and
`@libsql/client` ignoring `syncUrl` beside a remote url (a fact about
the client).

Readings at `7052b9111`. That head is the round's two commits plus a
true merge of `origin/main` `8cdbe0c6e`, which regenerated `registry.ts`
for its own new entry; the registry is current, with 308 semantic
entries.

- **CI on `2242ad513` before the push:** 35 check runs, 32 success, 3
skipped, 0 failed. Test Core 1/6, 3/6 and 5/6 and Type Check · workspace
all concluded success.
- **Tests:** `@objectstack/driver-turso` 75 files · 2014 passed · 18
skipped, typecheck exit 0. `@objectstack/spec` `--project local` in 3
shards: 6025 + 5141 (1 todo) + 5473 passed, exit 0 on each. A first
shard-2 attempt was killed by the runner's own timeout and re-run. Spec
typecheck exit 0.
- **Gates:** `dispatch-gates --commands` finds 9 paths vs merge base
`8cdbe0c6e` and 91 commands, all run with exit codes recorded before any
pipe. `--ran` reads `91 derived famil(ies) accounted for — 89 run, 2
NOT-MEASURED (2 DERIVED from a recorded exit 3)`, 0 UNRUN; the two NOT
MEASURED are `check:dual-build-cjs-loads` and `check:type-check-debt`
(whole-workspace build; both green in CI at `2242ad513`). Also run: the
five roster families under touched directories, and spec
`check:generated` ("All 15 generated artifacts are up to date"), all
exit 0.
- **Other readings:** `check:driver-conformance` is unchanged at 50
covered, 0 DEBT. The narrowed eslint run finds 8 files, 0 errors. The
control-byte scan finds nothing.

---
_Generated by [Claude
Code](https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN)_

---------

Co-authored-by: Claude <noreply@anthropic.com>
veigajoao pushed a commit to veigajoao/objectstack that referenced this pull request Sep 29, 2026
…te -wal sidecar too, never waiting on another connection (objectstack-ai#20426) (objectstack-ai#20463)

Fixes objectstack-ai#20426
Clause-②: no

`reclaimSpace()` on better-sqlite3 now returns the freed bytes from the
`-wal` sidecar as well as the freelist, and it never waits on another
connection. Every size below is the database file plus its `-wal` file,
read from the file system while the driver is still open. Every freelist
and page count is read from a second connection. Measured head:
`effb34a8a` (the branch after merging `origin/main` at `b28550818`,
which carries PR objectstack-ai#20427).

## What was wrong

With PR objectstack-ai#20425, one `Database.exec('PRAGMA incremental_vacuum')` returns
the whole freelist in one transaction. In WAL mode, the file-backed
default, that transaction's dirty pages outgrow the page cache, so
SQLite spills them into the WAL before the commit truncates them away.
Nothing afterwards truncates the WAL, so the sidecar keeps its
high-water size until the last connection closes.

## What changed

- `packages/drivers/driver-sql/src/sql-driver.ts`: the better-sqlite3
arm of `SqlDriver.reclaimSpace` calls a module-local
`reclaimBetterSqlite3(connection)`. It is module-local, like
`formatDuplicateGroups`, because `SqlDriver`'s `.d.ts` carries its
non-public members and this helper is no entry point. The published
types are unchanged; the `.d.ts` gains one doc-comment sentence on
`reclaimSpace`. The helper:
1. reads `PRAGMA freelist_count`, and sends nothing more when it is `0`;
2. runs `PRAGMA incremental_vacuum(N)` in chunks, N being a quarter of
this connection's page cache (1,000 pages at better-sqlite3's default
`cache_size = -16000` and 4 KiB pages), with a `PASSIVE` checkpoint
after each chunk;
3. stops when the freelist is empty or a chunk frees nothing (an
`auto_vacuum = NONE` file never shrinks its freelist);
4. ends with one `PRAGMA wal_checkpoint(TRUNCATE)` under a busy timeout
of `0`, and puts the connection's own busy timeout back in a `finally`.
Every statement goes through the binding's `exec()` / `pragma()`, which
step to completion. The loop is synchronous, so nothing else runs on the
connection between chunks. Every other SQLite client stays on
`knex.raw`, as before.
- Tests in `driver-sql` and `driver-turso` (below), and
`.changeset/20426-reclaim-space-wal-sidecar.md`
(`@objectstack/driver-sql`: `patch`).
- `.changeset/20106-reclaim-space-full-freelist.md`: one paragraph
removed. It said the freed pages pass through the `-wal` file, "which
keeps its size until the last connection closes". This PR makes that
false, and that note is still pending release. **This keeps
`check-empty-changeset` red on purpose** — see "The one red gate" below.

## The dispatch's hypotheses

- **H1 — confirmed** on `origin/main` `8cdbe0c6e`, through `SqlDriver`
(25,754 free pages):

  | step | database file | `-wal` | freelist / pages |
  |:--|--:|--:|:--|
  | after the delete | 103,149,568 | 4,255,992 | 25,754 / 25,789 |
  | after `reclaimSpace()` (351 ms) | 16,384 | 94,430,432 | 0 / 4 |
  | after one more write | 16,384 | 94,430,432 | 0 / 4 |
  | after `disconnect()` | 16,384 | 0 | 0 / 4 |

The DELETE-journal control on the same tree: 105,631,744 → 16,384 while
open, with no `-wal` file.

- **H2 — re-measured on this tree, and the picked variant is a fourth
one.** Each variant ran on `SqlDriver`'s own pooled connection after the
real fill-and-delete path (25,754 free pages, chunk 1,000). The rows
show database file + `-wal` after the call, driver open. This is one run
per cell on a shared box, so read the ratios, not the absolute times.

| variant | no reader | reader in this process (read transaction open) |
reader in another process (open for 1.5 s) |
  |:--|:--|:--|:--|
| `exec` alone (PR objectstack-ai#20425) | 16,384 + 94,430,432 · 315 ms | 103,149,568
+ 94,430,432 · 715 ms | 103,149,568 + 94,430,432 · 273 ms |
| + `wal_checkpoint(TRUNCATE)` | 16,384 + 0 · 476 ms | 103,149,568 +
94,430,432, busy · **5,333 ms** | 16,384 + 0 · **1,526 ms** (waited out
the reader) |
| chunked + `PASSIVE` | 16,384 + 4,255,992 · 157 ms | 103,149,568 +
4,255,992 · 47 ms | 103,149,568 + 4,255,992 · 62 ms |
| **chunked + `PASSIVE` + `TRUNCATE` at busy timeout 0 (this PR)** |
**16,384 + 0** · 276 ms, 108 ms on a rerun | 103,149,568 + 4,255,992,
busy · 48 ms | 103,149,568 + 4,255,992, busy · 61 ms |

- The objectstack-ai#20106 reading of about 210 KB for chunked + `PASSIVE` does not
hold through `SqlDriver`. `PASSIVE` never shrinks the sidecar: it stays
at whatever high-water size the sweep's own deletes left (4,255,992
here). Only a `TRUNCATE` checkpoint returns it.
- A waiting `TRUNCATE` checkpoint blocks the whole process on this
synchronous binding, for up to the connection's busy timeout (5,000 ms;
knex's better-sqlite3 client passes no `timeout`, so it is always
better-sqlite3's default). The lifecycle sweep runs in the server
process, so the triage's never-wait direction holds.
- So this PR takes the triage's chunked, never-waiting variant, plus one
`TRUNCATE` checkpoint that cannot wait. It is the only row that both
returns the space with no reader and never waits with one.
- With a reader present, no variant can shrink the database file. The
chunked rows keep the pair at its size before the call (107,405,560).
The one-statement rows grow it to 197,580,000.

**The chunk size, and why.** A chunk that outgrows the page cache spills
its pages into the WAL, just as one statement does. Frames left in the
WAL by the call, with a reader pinning every frame so none is reused:

| chunk (pages) | 100 | 250 | 500 | 1,000 | 2,000 | 4,000 | 8,000 | one
statement |
  |:--|--:|--:|--:|--:|--:|--:|--:|--:|
| default cache (`-16000`) | 1,437 | 1,121 | 1,003 | 928 | 883 | 3,779 |
14,216 | 22,920 |
  | 2 MB cache (`-2000`) | | 1,121 | 4,554 | 15,491 | | | | |

- The spill starts where the chunk reaches the page cache: between 2,000
and 4,000 pages at the default (`PRAGMA cache_spill` reads 3,871), and
between 250 and 500 at `-2000`.
  - Below that point, larger chunks mean fewer commits and fewer frames.
- A fixed 1,000 would spill on a connection with a smaller cache or
larger pages. So N is derived from the connection's own `cache_size` and
`page_size`, and the quarter leaves room for the per-page overhead and
the b-tree pages each chunk rewrites. At the default that is 1,000 pages
(4 MB).

- **H3 — confirmed.** `resolveSqliteJournalMode()` answers `wal` for a
file-backed database unless configured otherwise, and the probe's second
connection reads `journal_mode = wal`. The DELETE-journal control is
unchanged by the fix. Before and after, the file shrinks while the
driver is open and no `-wal` file exists: 105,631,744 → 16,384, 216 ms
before and 127 ms after.

- **H4 — confirmed.** The local `TursoDriver` face uses knex's
`better-sqlite3` client, so it takes this arm through
`super.reclaimSpace()`. Its suite reached the method, but it read only
the freelist and the page count. It now has a WAL-size case. The remote
route is untouched.

- **H5 — nothing new is thrown, so the sweep logs nothing new.** Both
checkpoints report "busy" as a result row, not as an error. So a busy
checkpoint degrades to "vacuumed, not checkpointed": the call resolves,
the pages are off the freelist, and `LifecycleService.sweep()` lists the
datasource as reclaimed, as before.
- Their bytes leave the files at a later checkpoint: the next reclaim
with free pages, SQLite's auto-checkpoint at 1,000 frames, or the last
connection closing. The reader case of the new test measures the next
reclaim.
- What can still throw is unchanged. Another connection holding the
write lock (`BEGIN IMMEDIATE`) makes the vacuum statement itself wait
out the busy timeout and throw `SQLITE_BUSY`. Measured: `main` 5,021 ms
and this PR 5,014 ms, both freelist unchanged, busy timeout 5,000
afterwards.
- In that case the sweep logs its existing warning (`space reclaim on
datasource 'X' failed (database is locked)`) and does not list the
datasource.
- The busy-timeout swap comes after the loop, so a throw inside the loop
never reaches it.

## The fix through `SqlDriver`

Same 25,754-page fixture:

| condition | database file + `-wal` after the call | call | busy
timeout after |
|:--|:--|--:|--:|
| WAL, no reader (was 103,149,568 + 4,255,992) | 16,384 + 0 | 101 ms,
109 ms | 5,000 |
| DELETE journal | 16,384, no `-wal` | 127 ms | 5,000 |
| WAL, reader in this process | 103,149,568 + 4,255,992 (unchanged;
freelist 0) | 47 ms | 5,000 |
| WAL, reader in another process | 103,149,568 + 4,255,992 (unchanged;
freelist 0) | 66 ms | 5,000 |

## Tests

`driver-sql/src/sql-driver-sqlite-reclaim-space.test.ts`, 7 cases (4
before). Each size is the database file plus the `-wal` file, read while
the driver is open. Each freelist and page count is read from a second
connection.

- **WAL:** freelist 0, and `{ file: pages × 4096, wal: 0 }` while open
and again after close.
- **WAL with a reader holding a read transaction.** The reopened file
has no WAL, the cache is set to about 100 pages, and 600 pages are free.
The case asserts:
  - the call resolves in under half the busy timeout;
  - the busy timeout reads 5,000 afterwards;
  - freelist 0;
- WAL growth under a quarter of the freed bytes. Measured: 0.08 for this
PR, and 0.87 for both one statement and a fixed 1,000-page chunk.
- Once the reader commits, the next reclaim returns everything: `{ file:
pages × 4096, wal: 0 }`.
- **The `auto_vacuum = NONE` control:**
  - the call resolves, so the loop stopped;
  - freelist and pages are unchanged;
  - the database file equals pages × 4096;
  - file + `-wal` is no larger than before.
- **DELETE journal:** freelist 0, and `{ file: pages × 4096, wal: 0 }`
while open.
- **The empty-freelist control, for both journal modes:** nothing
changes, the sizes included.
- **Unchanged:** the pooled connection is handed back.

`driver-turso/src/turso-remote-inherited-members.test.ts`: new case
"local face: in WAL mode the freed bytes leave the -wal sidecar too,
while the driver is still open".

Suites on the merged head `effb34a8a`, all through
`scripts/pm/os-verify-lock.sh`, each exit code recorded:
- `pnpm --filter @objectstack/driver-sql test`: exit 0, 195 files passed
and 11 skipped; 3,179 tests passed and 178 skipped. The count before the
merge was 3,227; the merge brought in PR objectstack-ai#20427, which removed tests of
its own.
- `pnpm --filter @objectstack/driver-turso test`: exit 0, 74 files;
1,982 passed and 16 skipped.
- `typecheck` for `driver-sql` and `driver-turso`: exit 0 each. `tsc
--listFilesOnly` shows both changed test files are in each package's
program.

## Ablations

Every leg ran on the committed state through
`scripts/ablation-replace.mjs`. In each, the anchor went from 1 hit to
0, and the restore was proven blob-equal to HEAD with an empty `git diff
HEAD`. The `driver-sql` suite imports `./sql-driver.js` (source), so
those legs needed no build.

| leg | mutation | result |
|:--|:--|:--|
| A | final `TRUNCATE` checkpoint removed | 3 red: WAL `{16,384 +
1,334,912}` vs `{16,384 + 0}`; the reader case's follow-up `{16,384 +
296,672}`; the NONE control's pair grew 1,318,384 → 2,555,376. 4 green.
|
| B | one statement instead of chunks | 1 red: the reader case, WAL
growth 2,142,400 vs a bound of 618,496. 6 green. |
| C | fixed 1,000-page chunk instead of the derived one | 1 red: the
reader case, 2,142,400 vs 618,496. 6 green. |
| D | busy timeout not zeroed for the `TRUNCATE` | 1 red: the reader
case, elapsed 5,034.99 ms vs under 2,500. 6 green. |
| E | busy timeout not restored | 1 red: the reader case, busy timeout 0
vs 5,000. 6 green. |
| F | the no-progress stop removed | the NONE control hung in the
synchronous loop and was killed after 60 s (SIGKILL). |
| A, dist | leg A built into `driver-sql`'s `dist/`, which
`driver-turso` resolves | `ablation-dist-preflight` found the marker in
2 built files. `driver-turso`: 1 red (local face `{32,768 + 280,192}` vs
`{32,768 + 0}`), 82 green. After the restore and a rebuild, `--absent`
found the marker in none of the 6 built files, and the tree was clean. |

In the first B–E runs, the red reader case also timed out its cleanup
hook: the failed assertion left the reader's transaction open. The fixed
case rolls the transaction back first. A rerun of leg B went red in 91
ms with no hook timeout.

## Gates

- `node scripts/pm/dispatch-gates.mjs --commands --repo
objectstack-ai/objectstack` at `effb34a8a` derived 63 commands. All 63
ran, and every exit code was recorded before any pipe. 62 exited 0;
`check-empty-changeset --base origin/main` exited 1 (next section).
- `--ran`: 63 derived, 63 run, 0 NOT-MEASURED, 0 UNRUN.
- The `--ran` pass printed a STALE TREE warning: `origin/main` moved 6
commits after the merge, and `scripts/cross-package-test-inputs.mjs`
changed in that range. Of those 6 commits, only PR objectstack-ai#20447 touches a
driver: it changes the `driver-turso` constructor, and none of this PR's
files. CI reads the merge ref.
- `check:driver-conformance`: 50 covered, 0 DEBT, 0 exempt, both before
(`8cdbe0c6e`) and after (`effb34a8a`).
- `pnpm lint` is CI's run. The narrowed run: `eslint --no-inline-config
--format json` over the 3 changed `.ts` files reports 3 files, 0 errors
and 0 warnings. `ESLint.isPathIgnored` answers false for each, so all
three are in `pnpm lint`'s population. `eslint.config.mjs` sets no
`parserOptions.project` and no typed rule, so this diff cannot move the
verdict of an untouched file.

## The one red gate: `check-empty-changeset` (a deliberate correction,
for confirmation)

This PR edits `.changeset/20106-reclaim-space-full-freelist.md`, which
exists on the merge base. The gate refuses that by name, and its own
text sets out two classes. This is the **deliberate correction** class,
not a collision.

- The removed paragraph says the `-wal` file "keeps its size until the
last connection closes". After this PR it is truncated at the end of the
call unless another connection is reading.
- That note has not been released, so restoring it from the base would
publish the false sentence.
- The gate's prescription for this class is to leave it red and get the
correction confirmed on the PR. `skip-changeset` is not applied and must
not be: this PR publishes a `patch`.

**For the seat: please confirm, or choose the other route.** The other
route is to restore the 20106 file from the merge base. The gate then
goes green, but the release would carry that sentence beside this PR's
own changeset, which describes the new behaviour.

## Acceptance notes

- **The file surface is widened by one file.** The claim names
`.changeset/20426-*.md`, and this PR also edits
`.changeset/20106-reclaim-space-full-freelist.md` (one paragraph
removed). It is the same defect, a mechanical removal, a card that has
already landed, and the same changeset gate family.
- **Behind a long reader, the bytes wait.** When a reader holds a
snapshot during the call, the database file keeps its size until a later
checkpoint. `LifecycleService.sweep()` still lists the datasource as
reclaimed. The next sweep that deletes rows returns it, and SQLite's
auto-checkpoint or the last close returns it sooner. No producer is left
worse off than on `main`, where the same reader left 197,580,000 bytes
instead of 107,405,560.
- **Partial progress is possible.** Chunks commit one by one. Another
process can take the write lock between two chunks, and then the next
chunk waits up to the busy timeout and may throw with the earlier chunks
already committed. This was not measured. A one-statement vacuum waited
and threw the same way, all or nothing.
- **Blocking is shorter, not gone.** The call still blocks the event
loop while it runs: 101 to 276 ms at 25,754 pages on this shared box,
against 315 to 351 ms for PR objectstack-ai#20425's single statement.
- The remote `TursoDriver` route, `SqliteWasmDriver`, `LifecycleService`
and `packages/spec` are untouched.

---
_Generated by [Claude
Code](https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN)_

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation size/m tests tooling

Projects

None yet

2 participants