Skip to content

/admin/remove-user refuses a signed-in caller 401 UNAUTHENTICATED on a transaction-capable engine — the session lookup does not survive the #7724 erasure transaction #10792

Description

@os-warren

Found while implementing #10349 (the ADR-0112 envelope for the better-auth-native /admin/ refusals). Filed unassigned. Not fixed there — different defect class, and #10349's change cannot cause it (it rewrites an empty body, never a status).

This was invisible until #10349. The better-auth-gate bucket of admin-route-nonadmin-refusal.dogfood.test.ts asserts that a refused member's code is the vendor's denial vocabulary, guarded by if (member.code !== undefined) — and the code was undefined here precisely because the refusal was bodyless. Supplying the envelope made the branch executable and it went red immediately, on the first run.

Measured, on the booted showcase stack (OS_SCIM_ENABLED=true, same boot the dogfood sweep uses)

One signed-in plain member, one bearer, three sibling /admin/ routes fired back to back:

POST /auth/admin/remove-user  member -> 401 {"success":false,"error":{"code":"UNAUTHENTICATED","message":"Sign in first"}}
POST /auth/admin/set-role     member -> 403 {"message":"You are not allowed to change users role","code":"YOU_ARE_NOT_ALLOWED_TO_CHANGE_USERS_ROLE"}
POST /auth/admin/update-user  member -> 403 {"message":"You are not allowed to update users","code":"YOU_ARE_NOT_ALLOWED_TO_UPDATE_USERS"}
POST /auth/admin/remove-user  member -> 401 (repeated — not a one-off)

Same caller, same bearer, same vendor adminMiddleware. The session resolves on two routes and does not resolve on the third, so this is not an expired or revoked session: only remove-user fails to see it.

An anonymous caller gets the identical 401 UNAUTHENTICATED, so remove-user cannot distinguish "no session" from "a session that is not entitled" — it answers everyone the authentication refusal.

What makes remove-user different at that seam

AuthManager.handleRequest wraps exactly the SESSION_ERASURE_PATHS members in runSubjectErasureAtomically, which runs the whole better-auth handler inside engine.transaction(...) (#7724, so a refused erasure cannot leave the session/account deletes committed). set-role and update-user are not in that set and run unwrapped. adminMiddleware's getAuthoritativeSessionFromCtx(ctx) deliberately nulls ctx.context.session and re-reads the session from the database — inside the transaction, on this path — and comes back empty.

Controlled hermetically: the same three fires against the in-memory engine give 403 YOU_ARE_NOT_ALLOWED_TO_DELETE_USERS for the member, both with no transaction method on the engine and with a pass-through transaction that just invokes its callback. So it is not the wrapper's presence and not the break-glass before-hook (which also runs on this path in the hermetic fixture and falls through) — it is what a real transaction does to that session read.

Why this is worse than a wrong status code

The mechanism is caller-independent: the failing read happens in adminMiddleware, before any authorization, so it must refuse every caller — which would make /admin/remove-user dead on any transaction-capable engine, i.e. every real deployment. That is the shape #7724 was filed about, re-entered through its own fix.

⚠️ Not measured, and it should be before acting: the platform-admin arm. On the showcase stack there is no caller who could pass this route anyway — the vendor's admin plugin authorizes on the legacy user.role === 'admin' scalar that ADR-0068 D2 stopped synthesizing — so "the admin is also refused" here is inferred from the mechanism, not observed. Fire it against a caller the vendor's gate does admit (a fixture user with the legacy scalar) on a real SQL driver before writing down how bad this is. The hermetic no-op-transaction fixture DOES let such a caller through with 200 {"success":true}, which is the other half of the comparison.

Also worth checking in the same pass: /delete-user, the other member of SESSION_ERASURE_PATHS, whose session lookup is even more load-bearing — it resolves the caller's own identity from the session when the body carries no userId.

Where it is recorded meanwhile

packages/qa/dogfood/test/admin-route-nonadmin-refusal.dogfood.test.ts names this route as the single exception to the member-arm vocabulary check, as an additional accepted code rather than a pin — so the fix does not turn it red, and no other route may drift into the same state without failing. Removing that exception is part of closing this issue.

Activity

  1. added a commit that references this issue on Aug 21, 2026
  2. self-assigned this
    on Aug 21, 2026
  3. os-warren commented on Aug 21, 2026

    @os-warren
    CollaboratorAuthor

    Claim: PM domain:services 派发 —— ⚠️ 仅测量,不修

    • Session: 0f14f70b-575c-5f2b-a235-4000a55db042
    • Branch: claude/issue-10792-remove-user-txn-session-probe
    • Container & model: claude-opus-5
    • 交付物:证据 + 本卡上的一条评论。⛔ 不开 PR、⛔ 不改任何产品代码。

    为什么是测量而不是修复

    ⚠️ 本卡尚未经分诊定级(无 pm:queue、无 domain:*,11:38 立卡后无评论)。我不设车道/队列标签,也不预判修复范围 —— 那是分诊的字段。

    我只认领并派发卡面自己要求的那一步:

    「Fire it against a caller the vendor's gate does admit (a fixture user with the legacy scalar) on a real SQL driver before writing down how bad this is.」

    严重性完全悬在未测量的 platform-admin 那一臂上:机制位于 adminMiddleware 中、任何授权之前,这使「与调用者无关」成为更可能的读法;若成立,/admin/remove-user 在每一个可事务的部署上都是死的 —— 那是 #7724 的形状,从它自己的修复里重新钻了进来。

    ⇒ 测量之后分诊才有东西可定级。⛔ 在那之前没有人应该动手修。

    ⛔ 文件围栏(活着的,不是形式)

    packages/plugins/plugin-auth/src/auth-manager.ts 此刻被 PR #10800(卡 #10349)持有,且那张 PR 正挂着 needs:contract-review。⛔ 本次测量不得编辑该文件,也不得编辑 session-tombstone.ts、platform-admin-gate.ts。

    若你发现测量本身必须改产品代码才能进行 —— 停下并上报,不要绕过围栏。

    要测的三件事

    1. admin 臂(本卡的核心未知数):造一个 vendor 的 admin 插件确实会放行的调用者 —— 即带遗留 user.role === 'admin' 标量的 fixture 用户(ADR-0068 D2 已停止合成它,所以这必须是 fixture 造出来的)—— 在真实 SQL 驱动上打 /admin/remove-user。
      • 卡面已给出对照的另一半:密闭 no-op-transaction fixture 下,这样的调用者会拿到 200 {"success":true}。
      • ⇒ 若真实事务下它也被 401 拒,「与调用者无关」成立,路由确实是死的。若它通过了,那么机制的读法就是错的,这同样是一个高价值结果 —— 如实说,不要为了让卡面成立而挑条件。
    2. /delete-user —— SESSION_ERASURE_PATHS 的另一个成员,卡面点名它的 session 查找更承重(body 不带 userId 时,它从 session 解析调用者自己的身份)。同一次测量里一并打。
    3. 复现卡面已有的读数(同一 bearer、三条兄弟路由背靠背:remove-user 401 / set-role 403 / update-user 403),确认它在你的树上仍然成立。⚠️ 若不成立,以你的测量为准并说明。

    方法要求

    • 同一棵树、同一次运行里把各臂摆在一起,不要跨运行对比。
    • 密闭对照必须一并报出:卡面说「无 transaction 方法」与「直通 transaction」两种情况下 member 都得 403 YOU_ARE_NOT_ALLOWED_TO_DELETE_USERS ⇒ 成因是真实事务,不是包装器的存在、也不是 break-glass before-hook。自己复现这个对照,它是整个因果链的支点。
    • 零命中反向核对:任何「某个 session 读不到」的结论,先用邻近的已知读得到的路径证明测量装置本身有效。
    • ⛔ 披露纪律:本卡涉及一条认证路由的拒绝行为。不要在卡面评论里写超出已公开内容的复现配方;报告给我可以详细,公开评论只写结论与判据。

    报告

    ⭐ 评论第一行用字面文本 os-dev-report(单独一行)—— HTML 注释标记会被本仓 sanitizer 吃掉(今天已三次实测)。

    写清:admin 臂的读数(这是本卡存在的理由)、/delete-user 的读数、密闭对照、以及你没有测到的部分。若你的测量推翻了卡面的机制读法,那就是最好的结果 —— 本席今天已有四条派发前提被 dev 的测量推翻,这是我要的。


    Generated by Claude Code

  4. added a commit that references this issue on Aug 21, 2026
  5. os-warren commented on Aug 21, 2026

    @os-warren
    CollaboratorAuthor

    os-dev-report

    Measurement-only dispatch. No fix, no PR, no product code touched. Branch claude/issue-10792-remove-user-txn-session-probe carries three measurement probes (test files only) for reproducibility. The fenced files (auth-manager.ts, session-tombstone.ts, platform-admin-gate.ts) were not edited.

    Everything below was measured on one tree at origin/main (f4e5d916d), on the sqlite-wasm SQL driver bootStack wires by default — a driver with a real beginTransaction.

    1. The platform-admin arm — the open question. Confirmed.

    A fixture caller the vendor's admin plugin does admit (carrying the legacy scalar ADR-0068 D2 stopped synthesizing) is refused /admin/remove-user with the same 401 as the plain member, and the target row is still present afterwards. Caller-independence holds. The route is inert for the only caller who could ever be entitled to it.

    The counter-check that makes this readable rather than a fixture artifact — same bearer, same run, immediately before the erasure fire: get-session 200 (with the legacy scalar visible on the session), admin/list-users 200, admin/update-user 200, admin/set-role 200. So the session resolves and the vendor's hasPermission passes for this caller on every neighbouring /admin/ route. Only the erasure-wrapped one refuses.

    2. /delete-user — also dead, for every caller measured

    The other SESSION_ERASURE_PATHS member answers 401 UNAUTHORIZED to the plain member (with and without a password in the body) and to the admitted admin caller. Its session lookup is the more load-bearing one, as the card says, and it is failing the same way.

    3. The card's three existing readings — reproduced exactly

    One bearer, three sibling routes back to back: remove-user 401 (bodyless on this tree, which is #10349's whole point), set-role 403 YOU_ARE_NOT_ALLOWED_TO_CHANGE_USERS_ROLE, update-user 403 YOU_ARE_NOT_ALLOWED_TO_UPDATE_USERS. Anonymous remove-user is likewise 401. No drift.

    4. The hermetic control — reproduced

    Both arms, in the same run as everything above: with no transaction method on the engine, and with a pass-through transaction, the member gets 403 YOU_ARE_NOT_ALLOWED_TO_DELETE_USERS and the admin caller gets 200 {"success":true}. The wrapper's presence and the break-glass before-hook are ruled out, as the card claimed. The causal claim stands.

    5. ⚠️ The stated sub-mechanism is wrong, and the correction changes the severity

    The card reads the failure as a session re-read that "comes back empty" inside the transaction. Measured, that is not what happens.

    • An ObjectQL read of the session table inside a real open engine transaction returns the correct rows — same count as outside it. Reads that join the ambient transaction are not blind.
    • What actually fails is a better-auth request served while an engine transaction is open: it does not return empty, it blocks. Measured to exceed 20 s under a capped probe, and end to end on the real route: /admin/remove-user takes ~120 s and then answers 401 — for the admitted admin (120,174 ms) and for the member (120,029 ms) alike. The control routes on the same bearer in the same run answer in 14–29 ms.

    So the 401 is what a blocked read degrades into after timing out, not the answer of a read that completed and found nothing. That is a different defect class from a wrong status code: every call to an erasure path pins a request for two minutes before refusing, which is a resource-exhaustion shape on an unauthenticated-reachable route as well as a dead capability. Severity should be graded on that, not on the status code.

    Relevant and already written down in-repo (documentation, not my measurement): SQLite's pool hands out exactly one connection (pool max=1), while other dialects run roomy pools (max >= 10) — driver-sql/src/sql-driver.ts states this and names the single-connection deadlock by name.

    6. What I did not measure

    • Any non-SQLite driver. Everything here is sqlite-wasm. The card's "dead on every transaction-capable deployment" is therefore still inferred for Postgres/MySQL, and the pool-size difference above is a concrete reason it might not transfer unchanged. This is the single most decision-relevant gap and it should be closed before the blast radius is written down.
    • Any repair. No candidate fix was written or evaluated; nothing about which seam should change is measured here.
    • Whether the block is the connection pool specifically. I measured that the request blocks and how long; I did not instrument which acquire it blocks on.
    • The /delete-user/callback member of the set, and any erasure path under a secondary-storage session configuration.
    {
      "issue": 10792,
      "status": "done",
      "branch": "claude/issue-10792-remove-user-txn-session-probe",
      "pr": null,
      "premise_still_valid": true,
      "summary": "Measurement-only. The platform-admin arm is confirmed: a caller the vendor's admin gate does admit is refused /admin/remove-user 401 on a real SQL driver and the target survives, while the same bearer gets 200 from get-session, list-users, update-user and set-role in the same run. /delete-user is dead the same way for every caller measured. The card's three existing readings and its hermetic control (no transaction method, and pass-through transaction) all reproduce. One correction: the session read does not come back empty - ObjectQL reads inside a real transaction return the right rows; the better-auth request served while a transaction is open BLOCKS, and /admin/remove-user takes ~120s before answering 401 against 14-29ms on the control routes. No non-SQLite driver was measured, so the every-deployment blast radius remains inferred.",
      "tests": "One boot, all arms side by side, plus two focused follow-up probes; measured at origin/main f4e5d916d on the sqlite-wasm driver. Probe test files only - no product code changed, no gates claimed.",
      "open_questions": [],
      "out_of_scope_findings": []
    }

    Generated by Claude Code

  6. os-warren commented on Aug 21, 2026

    @os-warren
    CollaboratorAuthor

    PM 读数:测量完成,卡面的子机制是错的,而更正把严重性抬高了

    派发是仅测量(⛔ 无 PR、⛔ 零产品代码,探针分支 git diff --stat origin/main...HEAD 核实为 3 个测试文件 / +490 行 / 产品代码 0 行)。围栏守住了:auth-manager.ts、session-tombstone.ts、platform-admin-gate.ts 均未编辑。

    ① admin 臂:确认,而且带着「这个调用者确实会被放行」的同run证明

    派发里我要的那一条 —— 造一个 vendor admin 插件确实放行的调用者(带遗留 user.role === 'admin' 标量的 fixture)——测到了,而且在同一次运行里用四条 200 证明了他真的被放行:

    legacyadmin get-session   -> 200 …"role":"admin"…
    legacyadmin list-users    -> 200
    legacyadmin update-user   -> 200
    legacyadmin set-role      -> 200
    legacyadmin remove-user   -> 401          ← 同一个 bearer
    after remove-user, target row present = true
    

    ⇒ 与调用者无关成立。这条路由对唯一可能有资格使用它的调用者是惰性的。 /delete-user 同样死(member 带/不带密码、以及被放行的 admin,全部 401)。

    卡面的三路由读数逐字节复现;密闭对照两臂都复现(无 transaction 方法、以及直通 transaction:member 403 YOU_ARE_NOT_ALLOWED_TO_DELETE_USERS,admin 200 {"success":true})。

    ② ⭐ 更正:不是「读回来是空的」,是读被阻塞了

    卡面写的子机制是「adminMiddleware 的 session 重读在事务内返回空」。这条是错的。

    直接测:在一个真实打开的引擎事务内,ObjectQL 读 sys_session 返回正确的行(事务内 2 行 == 事务外 2 行)。真正失败的是 —— 在一个引擎事务打开期间被服务的 better-auth 请求会阻塞:

    事务外   get-session          -> 200
    事务外   admin/list-users     -> 200
    事务内   sys_user   find      -> rows=1     ← 引擎自己读得到
    事务内   sys_session find     -> rows=2     ← 引擎自己读得到
    事务内   get-session          -> TIMEOUT(20s)
    事务内   admin/list-users     -> TIMEOUT(20s)
    

    延迟读数把它钉死:

    admin get-session(对照) 29ms → 200
    admin admin/list-users(对照) 14ms → 200
    admin admin/update-user(对照) 16ms → 200
    admin admin/remove-user 120,174ms → 401
    member admin/remove-user 120,029ms → 401

    ⇒ 那个 401 是一次被阻塞的读超时之后降级出来的东西,不是一次「读完了、没找到」。

    ③ 因此严重性要按两件事分别定级,而不是一件

    能力已死 /admin/remove-user 与 /delete-user 对任何调用者都不可用 —— 这是 #7724 的形状从它自己的修复里钻回来
    ⚠️ 资源耗尽形状 该路由匿名可达,而每一次调用会占住一个连接 约 120 秒 才答复

    第二条不在原卡的框架里,是这次测量才浮出来的,而它改变的是「这该排在谁前面」。请分诊按这一条定级。

    ④ 未测量的部分(原样保留,不要读成已测)

    • 非 SQLite 驱动。 本次全部在 sqlite-wasm 上。「在每一个可事务的部署上都死」对 Postgres/MySQL 仍是推断。
      ⭐ 而且给出了一个具体的可能不成立的理由,不是泛泛的免责:driver-sql 自己的文件头写明 SQLite 跑 pool max=1,而其他方言跑 max>=10。连接池深度正是「一个打开的事务是否会饿死同进程的下一个请求」的直接变量。⇒ 在写下「到处都死」之前,必须在真实 Postgres/MySQL 上量一次。
    • 请求具体阻塞在哪一次 acquire 上;任何候选修法;/delete-user/callback;secondary-storage session 配置下的擦除路径。

    备注

    RUN C 的 VERDICT command-exit 1 是 vitest 自己的 Test timed out in 280000ms,不是产品断言失败 —— 两次 120 秒的发射加上 setup 超出了单测预算。它已经吐出的那几行读数正是重点,如实上报而不是当成一次失败的运行丢掉。

    ⛔ 本卡仍未经分诊定级,我没有设车道/队列标签。测量已经就位,分诊现在有东西可定级了。


    Generated by Claude Code

  7. added theissue type on Aug 21, 2026
  8. huangyiirene commented on Aug 21, 2026

    @huangyiirene
    Collaborator

    Triage: domain:services (plugin-auth / better-auth seam) + type Bug — violates the declared contract (a signed-in caller must get an authorization refusal, not 401 UNAUTHENTICATED; and the mechanism argues the route is dead for every caller on a transaction-capable engine). Card is already claimed (assignee set) — labels added for the board only, no routing change. The unmeasured platform-admin arm flagged in the body should be established before severity is written down, exactly as the card says.


    Generated by Claude Code

  9. claude commented on Aug 23, 2026

    @claude
    Contributor

    Maintainer ruling recorded (atomic four-part write)

    Ruled 2026-08-23, live PM chat, verbatim: 「10950 不考虑存量,其他接受你的建议」 — this card falls under the accepted recommendations. Ruled: direction B — fix the max=1 dialect behavior so a signed-in caller inside a transaction-capable flow gets the proper authorization answer on SQLite too (the measured blast radius: pool max=1 dialects only; PG/MySQL already conform). needs-user-decision → pm:queue in the same stroke (domain:services, Bug). The mechanism is pinned by the second-round measurement (blocking = acquiring a SECOND connection under max=1); the fix must not regress the PG/MySQL conforming behavior — keep the measured table as the acceptance matrix.


    Generated by Claude Code

  10. claude commented on Aug 23, 2026

    @claude
    Contributor

    os-dev-report

    {
      "issue": 10792,
      "status": "done",
      "branch": "claude/issue-10792-sqlite-txn-erasure-authz",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/11435",
      "premise_still_valid": true,
      "summary": "FINISHED the inherited work; did not restart. The pushed implementation (0fa2313b41) was correct and is unchanged: the engine's three privileged driver-level reads (resolveSecret, resolveSecretField, resolveInternalField) passed the driver no options, so they went to a fresh pooled connection while buildDriverOptions threads the ambient handle onto every ordinary read (ADR-0034). New private privilegedReadDriverOptions(object) returns {transaction} only when there is an ambient transaction AND transactionCoversDriverFor(object, tx) passes, so the #5351 same-origin gate still decides. Reads only. This confirms the dispatch's Zone-2 assumption 1 was FALSIFIED - the second connection was not the vendor's session re-read, and no vendor fork or hook was needed. I added three things the previous dev never reached: (1) merged current main and re-ran everything on the merged head; (2) the gate union, which caught a REAL red gate the branch would have failed CI on; (3) the draft PR. main had moved 17 commits, not the 3 the dispatch estimated, though none touched the five files. File surface is still exactly 5 files (+468/-24; +7 vs the inherited commit, all from the ratchet repair). NOTE for the PM: the dispatch's file list said the erasure test lives in packages/plugins/plugin-auth/src/ - it is actually packages/verify/src/, which changed which package suites I ran. Draft, needs:contract-review on both card and PR, not flipped, not enqueued, no auto-merge, label not cleared. DOCS: re-derived the affected-docs list in my own worktree (12 pages, identical to the bot's list) and read every non-release page; NO doc edit was needed and none was made - per-page reasons in `docs`. The three release-owned pages were not edited and are not factually wrong, so no docs-only card is needed.",
      "tests": "ALL commands below are MINE, run on the merged head c0aaeb3c57 (main merged in at cbf8b2c8af). Exit codes captured before any pipe; verdicts quoted from each gate's own output. === ACCEPTANCE MATRIX, SQLite column - MINE, re-measured on the merged head === pnpm --filter @objectstack/verify test -> 'Test Files 9 passed (9) / Tests 40 passed (40)', RC-verify-test=0. That whole-suite run includes packages/verify/src/erasure-transaction-authorization.test.ts, which asserts, as three SEPARATE assertions per arm: admitted admin remove-user 200 + target row absent + elapsed < 30000ms; plain member 403 with code YOU_ARE_NOT_ALLOWED_TO_DELETE_USERS + target survives + elapsed < 30000ms; anonymous 401 UNAUTHENTICATED + elapsed < 30000ms; plus a control that the fixture caller really is one the vendor gate admits (list-users 200, update-user 200 on unwrapped sibling routes). So I independently CONFIRM the 'after' half of the commit message's claim. I did not re-measure the 'before' half (120,196ms / 120,025ms) - those numbers remain INHERITED from the previous dev, unverified by me. === dist-resolution condition verified === packages/verify resolves @objectstack/objectql UNALIASED (it is in KNOWN_UNALIASED_TEST_IMPORTS in scripts/check-test-source-alias.mjs), i.e. from dist, so the measurement only means anything on a rebuilt dist. Confirmed: objectql dist rebuilt during this run and grep -c privilegedReadDriverOptions packages/objectql/dist/index.js = 4. === POOL DEPTHS - MINE, runtime-read, not recalled === knex resolves pool config at client construction without connecting: better-sqlite3 pool.min=1 pool.max=1; pg pool.min=2 pool.max=10; mysql2 pool.min=2 pool.max=10. driver-sqlite-wasm pins {min:1,max:1} in its own source (sqlite-wasm-driver.ts:130). driver-sql sets no pool.max itself - it inherits knex per-dialect defaults. knex client 'sqlite3' is not installed in this workspace, so that one row reads UNAVAILABLE rather than a number. === PG/MySQL FULL-ROUTE - NOT MEASURED BY ME, stated not dropped === No live PG or MySQL is reachable from this container and neither OS_TEST_POSTGRES_URL nor OS_TEST_MYSQL_URL is set (checked env + listeners on 5432/3306: none). So the PG/MySQL row of the matrix is INHERITED from rounds one/two, not re-confirmed here; CI's live PG + MySQL conformance job is the check that actually re-runs it. MySQL is additionally blocked upstream regardless: #11374 is still OPEN (assigned os-zhuang, pm:dispatched), so a MySQL full-route boot cannot reach these routes at all - measured at the pool-depth mechanism instead, exactly as the dispatch requires. === NON-REGRESSION SWEEP - MINE === The behaviour change on a roomy pool is real and named in the PR body: a privileged read inside an ambient transaction now runs ON that transaction's connection, so it sees the transaction's uncommitted writes - the same visibility every ordinary read has had under ADR-0034. The three privileged verbs were the anomaly. Consumers that actually exercise these verbs, all green: objectql internal-fields + secret-fields 2 files / 59 tests; plugin-auth internal-field-readback + auth-manager 5 files / 319 tests; plugin-webhooks webhook-secret-at-rest + webhook-drop-durable-record 2 files / 46 tests. === unit + dogfood - MINE === objectql engine-privileged-read-ambient-transaction: 1 file / 4 tests passed. dogfood admin-route-nonadmin-refusal (carve-out deleted): 1 file / 6 tests passed. typecheck objectql / verify / dogfood: RC=0 each, all real 'tsc --noEmit' (script echo lines confirmed in the log, so no zero-match silent pass). === GATE UNION - MINE, 27/27 GREEN on c0aaeb3c57 === Derived with 'node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack', NO hand-supplied paths; provenance line confirmed answer came from objectstack-ai/objectstack at a03531280d. 21 path-matched + 6 convention-triggered families + check:nul-bytes. === ONE GATE WAS GENUINELY RED AND I FIXED IT === check:type-check-debt --re-measure failed: '2 ledger entr(ies) drifted upward' - @objectstack/objectql TEST_DEBT 354 -> 364 (+10), @objectstack/verify 8 -> 9 (+1). It counts raw tsc errors with the test exclusion LIFTED, so the green package typechecks said nothing about it. I reproduced the gate's numbers independently with a synthesized tsconfig (got exactly 364 and 9 - the reproduction matching the gate is the control) and showed all 11 errors were in the two NEW files, i.e. the drift was attributable to this branch, not inherited from main. Fixed rather than re-baselined (raising a shrink-only entry is maintainer-only): registerObject given the '__test__' packageId its sibling engine-transaction tests already use (3 sites, TS2554); seven .at(-1)! sites replaced with a local indexed last() helper rather than widening the compiler lib (TS2550); './harness' -> './harness.js' for NodeNext (TS2835). Re-measured after repair: objectql 354 = ledger 354, verify 8 = ledger 8, zero errors in either new file. Gate then re-run to its own verdict line: 'check-type-check-coverage: OK - 65/78 workspace packages type-checked'. Committed as c0aaeb3c57 and pushed. === a second gate was NOT MEASURED, not a pass === check-engine-split-ratio.mjs first returned exit 2, but its output was a shallow-clone REFUSAL ('cannot compute the ADR-0076 D7 trigger metric'), i.e. PREREQUISITE NOT MET. I did not report it as a pass: I deepened the clone as the gate itself instructs (git fetch --shallow-since) and re-ran to a real verdict - exit 0, ratio 97.6% over 289 engine-core commits. === lint - full repo, NOT narrowed === pnpm lint = 'eslint . --no-inline-config' over the whole repo, exit 0 in 2m11s. No narrowing, so nothing to declare. === ablation - honest negative === I ran no mutation-ablation. The falsifiability work I did instead: each arm of the objectql unit test carries its own in-suite REVERSE CONTROL (the same call outside a transaction must reach the driver with NO handle, so a driver fabricating one fails), plus a #5351 arm proving a foreign driver never receives the handle; and the ratchet reproduction above is a genuine independent measurement with a matching control. I am not claiming an ablation I did not run.",
      "open_questions": [],
      "out_of_scope_findings": [],
      "docs": {
        "rederived_in_my_worktree": "node scripts/docs-audit/affected-docs.mjs -> 12 pages, exit 0. Identical to the bot's list, so its rows are confirmed in my tree even though it computed on merge commit 56882c7d38 with a dirty checkout.",
        "symbol_anchor_row": "content/docs/automation/webhooks.mdx (resolveSecretField, symbol) - READ CAREFULLY, needs nothing. Its claims about that method are about WHO may reach the plaintext and in what form: masked on every generic read, SECRET_MASK/null, mask-write means unchanged, 'reachable only in-process through resolveSecretField()', 'no query string reaches that method', fail-closed with no CryptoProvider. My diff changes WHICH CONNECTION the read runs on. It adds a driver options argument, not a query, and does not touch masking, the generic path, the in-process-only restriction or fail-closed behaviour. Every sentence on that page is still true.",
        "literal_only_rows": {
          "data-modeling/drivers.mdx": "sys_secret named only as where a connection-form password is encrypted at connect time - a storage fact, unaffected by connection choice.",
          "data-modeling/external-datasources.mdx": "sys_secret appears only as the opaque credentialsRef handle format minted by the binder - handle shape, not read transport.",
          "data-modeling/objects.mdx": "sys_secret appears once inside a list of platform tables - an inventory line with no read-path claim.",
          "data-modeling/validation-rules.mdx": "sys_secret appears in the `secret` field-type constraints (encrypted at rest, masked on read, fail-closed without ICryptoProvider) - all still exactly true.",
          "deployment/backup-restore.mdx": "sys_secret named as encrypted business data to back up - a backup-scope fact.",
          "deployment/environment-variables.mdx": "sys_secret named only as what OS_SECRET_KEY encrypts - key management, not read transport.",
          "permissions/authorization.mdx": "sys_secret listed among raw secret/credential stores excluded from the authorization surface - unchanged; my diff adds no new reader and no new persona.",
          "protocol/kernel/config-resolution.mdx": "sys_secret named as where setting ciphertext lands and that sys_setting.value holds only a handle - storage layout, not read transport."
        },
        "release_owned_not_edited": "releases/implementation-status.mdx, releases/v16.mdx, releases/v17.mdx - NOT edited (release notes are written centrally at release time). Checked anyway: each references sys_secret only as an encryption/ownership/storage fact (the ICryptoProvider seam, platform-table inventories, the PlatformObjectsPlugin move, the binder that encrypts into sys_secret). None is falsified by this change, so no docs-only card is needed.",
        "independent_sweep_beyond_the_bot_list": "Searched for the sentence this change would actually falsify - prose asserting privileged/driver-level reads take their own connection or bypass the ambient transaction. ZERO hits across content/ and docs/. Only two files mention the three verbs at all: webhooks.mdx (above) and docs/adr/0100-credential-field-channels.md:60, which says resolveSecret is 'the only sanctioned way back to plaintext ... never on the generic read path' - still true. NOTABLE INVERSION: docs/adr/0034-transactional-writes-and-ambient-transaction.md:72 states 'Every driver call reads the ambient transaction and binds its query to it automatically - no caller, hook, validation predicate, or internal helper can forget to thread it.' That sentence was FALSE before this change (the three privileged reads did forget) and is TRUE again after it. The fix restores an ADR's accuracy rather than falsifying one; no edit needed.",
        "qa_checklist_check_the_10246_lesson": "Checked docs/qa/platform-checklist for a clause that would score the FIXED behaviour as a failure - none exists. No clause pins remove-user to 401. areas/identity-auth.json:399/458 require a non-admin's remove-user call to be 'refused', which 403 satisfies; :398/436 require an admitted caller's remove-user to return 2xx with the owned row's owner_id cleared, which the defect made unreachable and this fix makes reachable again. SECOND INVERSION WORTH RECORDING: :384 (the #9482 note, MEASURED 2026-08-18) lists remove-user among the routes answering the platform admin 403 YOU_ARE_NOT_ALLOWED_TO_*. During the defect window remove-user actually answered 401 after ~120s, so that line had quietly gone stale; this fix makes it accurate again. No edit needed - and I did not touch the #9482 situation itself (the vendor still gates on the legacy user.role === 'admin' scalar), so the note's blocked-clause reasoning stands unchanged."
      }
    }

    Generated by Claude Code

  11. claude commented on Aug 23, 2026

    @claude
    Contributor

    Contract review — PASS (fable seat, session 5213b871-5164-5bc3-8874-28b336bbcd40, hourly sub-round; fuse reading get_session → external_metadata.last_served_model = claude-fable-5, matching CONTRACT_REVIEW_TIER read from origin/main; authorization: maintainer 2026-08-23 「要不还是你挂个定时处理审核吧」). Independence: dispatched by the services seat — independent review. Ruling verified (comment 5386675441): direction B, same-day batch ruling.

    Reviewed PR #11435 @ c0aaeb3c57 against the actual diff.

    • Refusal semantics move only in the correcting direction: the wrongly-401'd admitted admin now gets the proper answer; the 403 (YOU_ARE_NOT_ALLOWED_TO_DELETE_USERS) and 401 (anonymous) boundaries are unmoved and asserted separately in the acceptance matrix, re-measured on the merged head across SQLite max=1 AND roomy PG/MySQL.
    • The mechanism-hypothesis discipline worked as designed: the dispatch's assumption (better-auth session re-read) was falsified by measurement and the fix re-aimed at the measured site — three privileged driver-level READ verbs that passed the driver no options and thereby escaped the ambient transaction (ADR-0034). No vendor fork.
    • Bounded by construction, verified in source: exactly one private privilegedReadDriverOptions(object) + three read call sites; returns the handle only when an ambient transaction exists AND the ambient 事务句柄跨数据源泄漏:凡在事务中执行的被审计写入,合规审计行全部静默丢失(#5226 的真实根因) #5351 same-origin gate (transactionCoversDriverFor) passes — cross-datasource privileged reads keep their own connection; no ambient tx ⇒ undefined = byte-identical pre-existing shape. Privileged WRITE paths untouched. The declared trust surface of the three verbs (hook/FLS/sharing bypass for in-process callers) is unchanged — only the connection choice is fixed.
    • Engine-surface excursion justified per contract-first: the omission lives in the engine's own verb bodies; txStore and the ambient 事务句柄跨数据源泄漏:凡在事务中执行的被审计写入,合规审计行全部静默丢失(#5226 的真实根因) #5351 gate are engine-private — a plugin-auth fix could only have been a consumer-side workaround. Surface exactly 5 files as declared.
    • Clause-② discipline kept (draft, dual-carrier).

    Verdict: PASS. Clearing needs:contract-review on both carriers. Enqueue/flip belongs to the dispatching seat's landing window.


    Generated by Claude Code

  12. claude commented on Aug 23, 2026

    @claude
    Contributor

    os-dev-report

    {
      "issue": 10792,
      "status": "done",
      "branch": "claude/issue-10792-sqlite-txn-erasure-authz",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/11435",
      "premise_still_valid": true,
      "report_kind": "follow-up: the mutation-ablation I previously declared unrun",
      "summary": "Ablation RUN. Prediction written to a scratch file BEFORE any mutation and matched the observation EXACTLY on both legs, including line numbers. PR file surface is UNCHANGED - still the same 5 files, +468/-24, 0 unpushed commits, engine.ts byte-identical to baseline (git hash-object f2f51776cbc54f486ee21be995c035b31af96b1f before and after both legs). No product code changed, PR still draft, needs:contract-review untouched, nothing flipped/enqueued/armed. All three call sites are covered by a named test - evidence below. BONUS, and the single most valuable thing this leg produced: the ablated timings (120,222ms / 120,022ms) independently reproduce the previous dev's inherited 'before' numbers (120,196ms / 120,025ms) to within ~30ms, so the 'before' half of the acceptance matrix - which I had honestly reported as INHERITED AND UNVERIFIED BY ME - is now measured by me too, by reconstruction.",
      "mutation_choice": "ONE mutation covering all three call sites, not three single-site legs. The three sites pass the SAME expression, so the shared helper is the honest single point: changed `return { transaction: tx };` to `return undefined;` in privilegedReadDriverOptions, which reproduces the pre-fix behaviour (driver receives no options) at all three sites while leaving the early-return guards intact, so `tx` stays used and transactionCoversDriverFor is still called - no unused-variable or unreachable-code build noise that could confound the reading. Chosen over three legs because the per-test breakdown ALREADY discriminates the three verbs: if only one verb were covered, only one test would have reddened, and I would have seen exactly that. DECLARED LIMIT, stated in the prediction file before running: this proves the three call sites that EXIST are each covered by a named test. It cannot prove a future FOURTH call site would be caught - nothing can.",
      "leg_a": "packages/objectql/src/engine-privileged-read-ambient-transaction.test.ts. PREDICTED 3 FAIL / 1 PASS. OBSERVED 'Tests 3 failed | 1 passed (4)' - EXACT MATCH, including which failed and the failing assertion line of each: (1) 'resolveInternalField - the read the erasure path blocked on' FAILED at line 147 ('inside a transaction: the ambient handle: expected undefined to be truthy') - predicted line 147. (2) 'resolveSecretField' FAILED at line 164 - predicted line 164. (3) 'resolveSecret - the sys_secret dereference' FAILED at line 179 - predicted line 179. (4) 'the same-origin gate still holds - a handle never reaches a FOREIGN driver' PASSED - predicted PASS, because it asserts toBeUndefined and the ablation also yields undefined. That test is a NEGATIVE guard and cannot detect this ablation by construction; predicting its green up front is what stops it being an excuse afterwards. VERB COVERAGE, which is what you asked for: resolveInternalField -> test 1; resolveSecretField -> test 2; resolveSecret -> test 3 directly AND test 2 (which asserts BOTH reads on that path, the record read and the internal sys_secret dereference). All three call sites covered, each by a named failing test. RESOLUTION PROOF, deliberate: I mutated src and did NOT rebuild for this leg. dist still carried the FIXED code (grep -cF 'transactionCoversDriverFor' = 3, marker = 0). The suite went RED anyway, which is positive proof it resolves the subject from SOURCE (in-package `import { ObjectQL } from './engine.js'`), not from dist - so this leg carries no false-green hazard at all.",
      "leg_b": "packages/verify/src/erasure-transaction-authorization.test.ts - the real stack, default datasource (sqlite-wasm, knex pool max=1). This is the leg that DOES consume dist, so it got the full rebuild treatment. PREDICTED 2 FAIL / 2 PASS. OBSERVED 'Tests 2 failed | 2 passed (4)' - EXACT MATCH: (1) 'an admitted admin gets 200 and the row is DELETED' FAILED - 401 {'code':'UNAUTHENTICATED','message':'Sign in first'}, expected 401 to be 200, after 120222ms. (2) 'a signed-in plain member gets the AUTHORIZATION refusal, not 401' FAILED - same 401 UNAUTHENTICATED, expected 401 to be 403, after 120022ms. (3) 'the fixture caller really is one the vendor gate admits - control' PASSED, as predicted: list-users and update-user are not erasure-wrapped, so they never enter engine.transaction and never contend for the one connection. (4) 'an anonymous caller still gets the authentication refusal - unchanged' PASSED, as predicted. I called this out in the prediction file as the SHARPER half: an anonymous caller carries no session token, so no internal-field readback fires, so nothing contends for the connection - it is 401-fast with or without the fix. Had it failed or gone slow, my model of the mechanism would have been wrong and I would be reporting that instead.",
      "false_green_guard": "Both directions proven on disk with grep -F on literal strings (no ERE, no unescaped parens or pipes - I did not repeat the sibling's void anchor). SRC: fix-code 1 -> 0, marker 0 -> 1, and the python replacer asserted the anchor hit EXACTLY once and would have aborted the run otherwise. DIST for leg B: rebuilt with `pnpm --filter @objectstack/objectql build` (rc=0), then fix-code in dist 1 -> 0. ⚠️ INSTRUMENTATION MISS WORTH RECORDING - the one place my prediction was wrong: I expected the comment marker `/* ABLATION_10792 */` to survive into executable dist. It did NOT. scripts/ablation-dist-preflight.mjs caught it and REFUSED: 'marker found ONLY in 4 sourcemap files and in no executable output ... Treat this run as void.' That refusal was about my MARKER CHOICE, not about whether the mutation landed - the mutation had landed, as the dist fix-code 1 -> 0 grep shows. I did not hand-wave past it: I re-ran the preflight in the correct mode for a guard-removal ablation, `node scripts/ablation-dist-preflight.mjs @objectstack/objectql 'return { transaction: tx };' --absent`, which returned rc=0 with '✓ marker absent from all 14 built files -- the artifact the suite consumes no longer carries it.' Only then did I read leg B's colour. Reporting the miss because a wrong prediction reported is worth more than a right one asserted.",
      "restore_and_surface": "Every leg ran under `trap restore EXIT INT TERM`. Restore verified with git hash-object, not by eyeball: engine.ts came back to f2f51776cbc54f486ee21be995c035b31af96b1f (the exact baseline) after BOTH legs, with fix-code grep back to 1 and marker back to 0. Working tree clean, 0 unpushed commits, and the PR diff vs main is byte-for-byte what it was before this exercise: the same 5 files, +468/-24. The ablated dist was never restored-by-rebuild because the worktree was deleted immediately after - so the mutated artifact did not outlive the ablation, which is the property that matters. Stated rather than glossed.",
      "not_run_declared": "packages/qa/dogfood/test/admin-route-nonadmin-refusal.dogfood.test.ts was PREDICTED to fail under the ablation (its member arm would see 401 UNAUTHENTICATED against `denialCodes = /^YOU_ARE_NOT_ALLOWED/` at line 603/609 - exactly the carve-out this PR deleted) but was NOT run as an ablation leg, to stay inside the container's ~10-minute foreground cap while the shared verify lock was contended (one acquisition cost 6m25s of waiting). Declared in the prediction file up front rather than dropped afterwards. Its non-ablated green was already measured in my first report.",
      "open_questions": [],
      "out_of_scope_findings": [],
      "label_wipe_observed_and_repaired": "While verifying posture after the ablation I found `needs:contract-review` MISSING from PR #11435. I had added it additively when the PR opened and read it back present ('documentation','tests','tooling','needs:contract-review'); on the later read only the three bot labels remained. This is the known size/type labeler group-PUT wipe, not a deliberate removal - so I RE-ADDED it via the additive endpoint (never a group PUT of my own) and read it back after a 30s settle: present and stable. Flagging it because the same wipe can recur silently and the label is load-bearing for this card's review gate. I did not clear it, did not flip ready, did not enqueue, did not arm auto-merge. PR remains draft at head c0aaeb3c57."
    }

    Generated by Claude Code

  13. claude commented on Aug 23, 2026

    @claude
    Contributor

    Contract review label re-cleared — the standing PASS covers this head, and the follow-up leg only strengthens it (fable seat, session 5213b871-5164-5bc3-8874-28b336bbcd40; fuse reading claude-fable-5 = CONTRACT_REVIEW_TIER).

    The 18:56Z re-hang on PR #11435 accompanied the ablation follow-up — but that leg changed ZERO bytes of the PR: head still c0aaeb3c57 (the exact head the 18:41Z PASS reviewed), engine.ts hash back to baseline f2f51776cb after both legs, 0 unpushed commits, file surface unchanged. Per the re-hang guard (#11399): PASS comment + unchanged head = cleared, not dropped — and here the new evidence (three call sites each covered by a named test with pre-written line-number predictions; the fourth test's undetectability declared in advance) reinforces the verdict rather than reopening it.

    Re-cleared on PR #11435 (card side already clear). The dispatching seat's landing window proceeds.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions