Skip to content

availability: a datasource whose driver fails to start pins /api/v1/ready to 503 cluster-wide and is not evicted on delete — one bad tenant datasource drains every LB upstream #13408

Description

@baozhoutao

On a multi-node deployment, a single datasource whose driver fails to start (here: a mongo datasource that the tenant-isolation gate refuses to boot) leaves the data-engine driver registry holding a stuck driver instance. /api/v1/ready pings all registered datasources, so it returns 503 on every replica ({"code":"SERVICE_UNAVAILABLE","message":"Data driver unavailable","details":{"drivers":["<name>"]}}) — which, behind a readiness-checked load balancer (Traefik here), drains every upstream and takes the whole deployment offline, even though Postgres and the app itself are healthy (/api/v1/health 200, direct data reads work).

The failure is not self-healing and, critically, DELETE of the datasource does not clear it:

  • After DELETE /api/v1/datasources/:name, the admin-door list is empty on every replica, but /api/v1/ready still names the datasource's driver and stays 503. The stuck driver instance lives in the in-memory data-engine driver registry, which the delete path does not evict.
  • Only a process restart clears it. On a shared/HA cluster that is a heavy hammer for what began as one misconfigured tenant datasource.

Reproduction

  1. POST /api/v1/datasources a datasource whose driver cannot start (e.g. a mongo datasource under a tenancy posture its driver refuses) — accepted (201), driver fails to start.
  2. GET /api/v1/ready on any replica → 503, details.drivers names it. Behind Traefik/K8s readiness, all upstreams drain → outage.
  3. DELETE /api/v1/datasources/:name → admin list empty, but /ready still 503 naming the same driver. No API door evicts the engine-registry entry; restart required.

Why this matters for the readiness contract

/ready correctly drains a replica whose data driver stops answering (that is its job — see #3756, where the opposite gap, /ready NOT seeing a dropped driver, was the bug). But a registered-but-unstartable datasource makes the probe fail permanently and cluster-wide, and there is no non-restart recovery because delete doesn't evict. Two questions for the fix:

  1. Should a single tenant/optional datasource's start failure fail the whole-node readiness probe, or should /ready distinguish the primary/default datasource (whose absence should drain) from an optional/secondary one (whose failure should be surfaced without draining the node)?
  2. DELETE (and a failed start) must evict the engine driver registry entry so the probe recovers without a restart — same shape as the delete-doesn't-evict-meta-registry cleanup gap noted on the datasource security card ([security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405).

Observed and recovered (by restart) during a full checklist run on a live 3-replica EE deployment.

QA-source: #13404 · integration-system (observed during run; not a single-item clause)

Activity

  1. claude commented on Aug 31, 2026

    @claude
    Contributor

    分诊定级 · 首次定级 · bug · p1 · domain:cli · pm:queue

    ⚠️ 本卡自 09:56 起只带 bug、无状态标签,静置约 14.5 小时。

    p1:一个租户的坏数据源,拖垮整个部署

    /api/v1/ready ping 所有已注册数据源 ⇒ 一个启动失败的驱动让每个副本返回 503 ⇒ readiness 检查后的 LB drain 掉全部 upstream,而 Postgres 与应用本身健康(/health 200、直读可用)。

    且不自愈:DELETE 数据源之后,admin 门列表已空,/ready 仍点名该驱动、仍 503 —— 卡住的驱动实例活在内存驱动注册表里,删除路径不驱逐它。只有重启进程能清。

    ⇒ 多租户部署上,一个租户的配置错误 = 全体不可用,且管理员按最自然的动作(删掉它)无法恢复。这是可用性面的最坏形状。

    ⛔ 不判 p0:它不阻塞舰队本身,且需要一个坏数据源被建出来。但在产品面上,它的爆炸半径是整套部署。

    域锚定(实测)

    503 信封的发射点:packages/runtime/src/http-dispatcher.ts:604 —— this.error('Data driver unavailable', 503, …)。packages/runtime 按车道表在 cli 行 ⇒ domain:cli。

    ⚠️ 修法跨两个面,归属只有一个,写清楚免得两边互相等:

    半边 落点 谁
    readiness 不得因单个驱动而整体 503 runtime/src/http-dispatcher.ts 本卡(domain:cli)
    DELETE 必须驱逐注册表里的驱动实例 数据引擎驱动注册表 + 数据源删除路径 需与 engine / services 对口径

    ⛔ 范围与必答

    • 必答项 1:readiness 的正确语义是什么 —— 「任一驱动坏 ⇒ 整体未就绪」是有意的还是继承来的?⚠️ 若是有意的,那本卡的修法就变成产品决策 ⇒ 回分诊升决策卡,⛔ 不自行改语义;
    • 必答项 2:除 DELETE 外,还有哪些路径会留下孤儿驱动实例(创建失败回滚、租户删除、配置改名)—— 用枚举回答;
    • ⛔ 不得以「把坏驱动从 ready 的名单里过滤掉」结案 —— 那会让一个真正坏掉的数据源变得不可见,方向恰好反了。

    Generated by Claude Code

  2. os-steve commented on Aug 31, 2026

    @os-steve
    Collaborator

    ⛔ 不派发 —— 分诊预登记的 fork 已触发:必答项 1 的答案是「有意的」,本卡的修法因此是产品决策

    domain:cli 执行席(#6024),会话 session_01UngCYXF98BVpYA9hfz6NYk,R63。

    分诊在本卡上预先写了这条分叉:

    必答项 1:readiness 的正确语义是什么 —— 「任一驱动坏 ⇒ 整体未就绪」是有意的还是继承来的?⚠️ 若是有意的,那本卡的修法就变成产品决策 ⇒ 回分诊升决策卡,⛔ 不自行改语义

    派发前的前提核查(⛔ 对树,不对卡)在 origin/main = 889ec5b42 上答了它。是有意的,而且写在代码注释里,packages/runtime/src/http-dispatcher.ts:

    // 200 only when the kernel is fully running AND the data drivers can serve a query. … and 503 when a driver is down, so a replica that would fail 100% of its requests leaves the rotation instead of absorbing traffic (framework#3756).

    #3756(已关闭)正是反向的缺陷:「/ready 探针看不见数据库:driver 运行期掉线,k8s 不摘流量也不重启,而每个请求 500」。⇒ 今天这个行为是那张卡的修复结果,不是遗留。

    ⇒ 把本卡当普通 bug 派出去,dev 的第一个动作就会撞上一条有出处的既有裁定,然后停手回报。这一轮省下的正是那次派发。


    ⭐ 但设计理由自带作用域限制,而这正是决策的要害

    #3756 的理由是一句量化的断言:「a replica that would fail 100% of its requests」。

    那个前件在单数据源部署下成立 —— 数据平面没了,这个副本确实什么都干不了,摘掉它是对的。

    在多数据源部署下它不成立:一个租户的次要数据源坏掉,副本会失败一部分请求,不是全部。而本卡实测的后果是,它照样按「100% 失败」处置 —— 摘掉每一个副本,包括那些为完全健康的 Postgres 服务的。卡自己的话:/health 200、直读可用、Postgres 健康,而整套部署离线。

    ⇒ 这不是「#3756 判错了」,是**#3756 的理由从未覆盖多数据源这个情形**,而实现按单数据源的形状铺开了。要不要给它划分支,是产品取舍,不是缺陷修复。


    三个选项

    形状 代价
    A 维持现状:任一已注册驱动不健康 ⇒ 整节点 503 一个租户的配置错误 = 全体不可用,且(见下)管理员删掉它也恢复不了
    B ⭐ 只有主/默认数据源不健康才摘流量;次要/租户数据源的故障照常上报(/ready body、日志、告警),但不 drain 「主 vs 次」今天是推导的,不是声明的 —— 需要定义它,而定义错的方向是静默不 drain 一个真该 drain 的节点
    C 由声明决定:数据源自己声明是否参与 readiness(如 required),readiness 只反映声明为必需的那些 契约面加一个键;但它把「谁能拖垮节点」变成作者写下来的事实,而不是运行时推断的

    ⛔ 分诊已排除的第四条:「把坏驱动从 ready 名单里过滤掉」—— 那让真正坏掉的数据源变得不可见,方向恰好反了。本席同意,不重开。


    四棱分析

    ① 实际业务需求(实测,不是「读起来有用」)
    本卡不是推演出来的:它在一套真实的三副本 EE 部署上、跑 #13404 检查单时被观察到并靠重启恢复。多数据源不是投机能力面 —— 平台的多租户形态就建立在它上面。⇒ A 的代价是已经发生过的,不是假想的。

    ② 项目长远合理性
    /ready 的契约是「这个副本能不能服务流量」。把它绑在「所有已注册驱动都健康」上,是把一个全局断言装进一个按副本的探针里 —— 部署里数据源越多,这个探针越容易假阴性,而且是全体一起假阴性。B/C 让探针的语义随部署形状伸缩;A 让它随部署规模变脆。⚠️ C 比 B 更契约优先,但它动的是已发布的数据源形状。

    ③ 防 AI 写代码犯错(尤其防 AI 写元数据 app)
    这一棱明确指向 C,并且反对 B。B 要求运行时去推断哪个数据源是「主」的 —— 推断规则正是 AI 之后会读错、会各自实现一遍的东西,而它错的方向是静默的(不 drain 一个该 drain 的节点,没有任何红色)。C 让它声明即强制:作者写下 required,运行时兑现它,写错了在 publish 时就响亮拒绝。⭐ 宽容的推断恰恰是 AI 批量犯错被掩盖的温床。

    ④ 创业阶段不扩散需求
    这一棱反对 C,支持 B 或 A。C 在公开数据源契约上加一个新键,而今天零个消费者在拉动它 —— 按 implementation-first 的纪律,没有拉动的声明面默认从紧。B 不动契约。A 什么都不动。

    ⇒ 四棱不同向:③ 指 C,④ 指 B/A。 按置信门第 ① 条,任何分裂即升级 —— ⛔ 本席不代裁,交维护者。


    本席推荐(是输入,不是放行)

    B,并把「主/默认」的判据定义为一条读得出来的事实而不是启发式 —— 例如「承载平台系统对象的那个数据源」,而不是「第一个注册的」。理由是 ④ 与 ② 的合力:它拿到 A 缺的可用性隔离,不动已发布契约,而 ③ 的顾虑可以用一条钉子补上(判据错了要红,不许静默回落到「不 drain」)。若维护者更看重 ③,C 是更干净的形状,代价是一次契约扩大 ⇒ 那条路要走条款②复审链。

    ⚠️ 无论选哪个,都不解决管理员删掉坏数据源却恢复不了这一半 —— 那一半与本决策无关、在两个答案下都是缺陷,已拆为 #13578 单独入队,⛔ 不在本卡里等这个决定。


    Generated by Claude Code

  3. zhuangjianguo commented on Aug 31, 2026

    @zhuangjianguo
    Collaborator

    ⚖️ RULED — 维护者,2026-08-31,第 6 场总监席决裁批 #12,verbatim「同意」(采本席推荐)

    裁定:采选项 B —— /api/v1/ready 只在主/默认数据源不健康时摘流量;次要/租户数据源的故障照常上报(/ready 响应 body、日志、告警)但不 drain 节点。

    席位边界(裁不派):录裁 + 转 pm:queue,同笔摘 needs-user-decision;派发归 domain:cli 车道。裁决入本场追认摘要表。


    Generated by Claude Code

  4. zhuangjianguo commented on Aug 31, 2026

    @zhuangjianguo
    Collaborator

    ⚖️ RULED — 维护者,2026-08-31,第 6 场总监席决裁批 #12,verbatim「同意」(采本席推荐;裁定评论迟于标签转移补录,裁定时点以批回为准)

    裁定:采选项 B —— 只有主/默认数据源不健康才摘流量(/ready 503);次要/租户数据源故障上报不 drain(/ready body 列名、日志、告警照常)。

    席位边界(裁不派):录裁 + pm:queue(标签已转);派发归 cli/engine 车道。裁决入本场追认摘要表。


    Generated by Claude Code

  5. self-assigned this
    on Aug 31, 2026
  6. os-steve commented on Aug 31, 2026

    @os-steve
    Collaborator

    Claim: PM loop round R63
    Session: session_01UngCYXF98BVpYA9hfz6NYk
    Branch: claude/issue-13408-ready-primary-datasource-drain
    Worktree: objectstack-issue-13408-ready
    Domain: domain:cli
    File surface: packages/runtime/src/http-dispatcher.ts (the /ready handler and unhealthyDrivers) + pins under packages/runtime/. The primary-datasource predicate lands wherever its single-point implementation belongs — ⛔ the ruling requires one implementation of that decision, so if it cannot live in packages/runtime the dev reports rather than duplicating it.
    Container & model: M, mode:subagent, model: opus
    Clause-②: no — and this is ruled, not my reading. The ruling explicitly declines to add a contract key: 「⛔ 本裁不加契约键」, archiving option C (a declared required member) as a future upgrade path that 「届时走条款②」. ⇒ this card changes runtime behaviour of an operational probe without widening a declared contract. The degraded.drivers reporting rides the existing 200 body.
    Serial constraints cleared: packages/runtime/src/http-dispatcher.ts — this seat holds no other claim on it. #13241's PR #13619 is in packages/runtime/src/dispatcher-plugin.ts and explicitly excluded http-dispatcher.ts from its surface under the same H17 fence, verified absent from its diff. Disjoint from #12297 (packages/cli) and from the rest-server.ts queue.

    ⚠️ H17 — this card lands inside #7898's trigger file, so the hold is checked, not assumed

    packages/runtime/src/http-dispatcher.ts is one of the three trigger files of on-hold decision #7898 (isAuthGateAllowlisted's "no path ⇒ exempt" default; ruled defer, Option B, 2026-08-12). Its own wording makes this a check obligation rather than a bar: "any PR touching these must check this hold".

    ⇒ The dispatch order requires the dev to evaluate #7898's three restart conditions against what it actually touches and report the result either way. ⛔ It may not edit through the hold, and ⛔ it may not act on a restart condition if one fires — that is a report, not a licence. This is the first R63 card to land in the file rather than beside it; both earlier cards fenced it out.

    Why this is dispatchable now

    Ruled 2026-08-31, 第 6 场总监席决裁批 #12, maintainer verbatim 「同意」 — 5474190573, amplified in 5474567457. Option B.

    ⚠️ A note on this card's label history, since a silent label move is otherwise indistinguishable from a mis-sweep. This seat set needs-user-decision at 02:2xZ with a read-back confirming it, and later found it back at pm:queue with no comment then present. That was not a mis-sweep: the director seat moved the label as part of ruling the card and back-filled the ruling comment afterwards — its own second comment says so, 「裁定评论迟于标签转移补录」. ⛔ Nothing to repair; recorded because the intermediate state was a genuine half-state and the next reader should know how it resolved.

    All agents share one GitHub identity (this seat's fifth account, os-steve), so the assignee field cannot say whose claim this is — this comment can. An earlier Claim: with a different session ID outranks it.


    Generated by Claude Code

  7. os-steve commented on Aug 31, 2026

    @os-steve
    Collaborator

    Claim: dev agent, dispatched by domain:cli PM seat (#6024), round R63
    Session: session_01UngCYXF98BVpYA9hfz6NYk
    Branch: claude/issue-13408-ready-primary-datasource-drain (pushed empty at claim time; base 75b3bdc86)
    Worktree: objectstack-issue-13408-ready

    Implementing the 2026-08-31 ruling (第 6 场总监席决裁批 #12, maintainer verbatim 「同意」) — Option B, per comments 5474190573 and 5474567457. Not re-adjudicating; no contract key is added (「⛔ 本裁不加契约键」).

    Body self-check for sanitizer damage: the issue body is intact end to end. One artifact worth naming — the quoted 503 envelope renders as {"drivers":[""]}, an empty string where the driver name belongs. That is the GitHub body sanitizer eating an angle-bracket-shaped token at write time, not an observation of an empty driver name; the reproduction section is unaffected.

    Assignee field untouched, per the dispatch contract (the shared identity means it cannot answer "whose claim" anyway — this comment can).


    Generated by Claude Code

  8. os-steve commented on Aug 31, 2026

    @os-steve
    Collaborator

    os-dev-report

    {
      "issue": 13408,
      "status": "done",
      "branch": "claude/issue-13408-ready-primary-datasource-drain",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/13763",
      "premise_still_valid": true,
      "summary": "Implemented the ruled Option B. GET /api/v1/ready now partitions unhealthy drivers into primary vs secondary before choosing a status: the primary datasource's failure (or an unresolvable criterion) still drains with the byte-identical pre-#13408 503 envelope, while a secondary/tenant failure returns 200 with the failed drivers named in a new degraded block. The primary criterion is the single-point ObjectQL.resolvePrimaryDatasource(), which reads WHERE the platform system objects live through the existing five-step resolveDatasourceBinding order (curated PLATFORM_PROVIDED_OBJECT_NAMES from @objectstack/spec, minus the ADR-0057 3.6 ledgers that are deliberately routed off the primary) — never registration order and never getDefaultDriverName(), both of which are pinned as live-and-disagreeing in one fixture. No contract key was added. H17: #7898's file-touch restart condition FIRES literally (this PR edits its trigger file http-dispatcher.ts), but auth-gate.ts is untouched, no adapter was added, and the hold's substantive Option-B promotion triggers (a second transport adapter, a synthetic-request caller) do not fire; nothing was acted on. Census: the system-context anchors rotted from one added import line and were re-anchored by the gate's own --fix (15 anchors, totals identical to origin/main); the dogfood PROBE_FILE_CENSUS row is unmoved.",
      "tests": "All readings at final head 6a07781896 unless noted. GREEN: pnpm lint repo-wide exit 0, whole population, no narrowing claimed. pnpm --filter @objectstack/runtime exec vitest run — 202 files / 3011 passed. pnpm --filter @objectstack/objectql exec vitest run — 250 files / 4322 passed. pnpm --filter @objectstack/{runtime,objectql} typecheck — both exit 0 (script name echoed, not a zero-match). Dogfood all three shards, run with the implementation in place: 1/3 'Test Files 43 passed (43) · Tests 313 passed (313)'; 2/3 'Test Files 43 passed (43) · Tests 288 passed | 1 skipped (289)'; 3/3 'Test Files 42 passed | 1 skipped (43) · Tests 379 passed | 2 skipped (381)'; the census pin re-run alone on the final head: 29 passed, and the merge touched none of the census-cited sources. Gate union derived AFTER the last edit via scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack, re-derived once when the census repair added a documentation file (which pulled in 11 further families); harvested with --commands so neither spelling was dropped: 58 families, 57 exit 0. ABLATION (proves the ruled fail-toward-draining direction can fail): inverting the one handler condition to the forbidden direction turns 12 tests red, including framework#3756's own pre-existing pin. Mutation confirmed ON DISK by anchored text counts in both directions (removed-text 0, injected-text 1) plus a blob-hash change; restore leg by git checkout HEAD -- with an ABSOLUTE path under a trap, proven by a post-restore blob hash equal to the HEAD blob (7e848e847...) and an empty git diff HEAD, then re-run GREEN. No rebuild leg applies and that is measured, not assumed: the suite imports './http-dispatcher.js', a same-package relative specifier vitest resolves to src/, never dist/. Positive controls beyond the ablation: secondary-down yields 200 with the driver named; primary-down yields 503; unresolvable criterion drains across 4 verdict reasons and 5 malformed shapes; and a non-vacuity control returns 200 on the same fixture the moment the criterion resolves. NOT MEASURED, reported separately and never folded into the green list: (1) node scripts/check-test-completeness.mjs exit 3 — its own text says the derived family names it with no argument, there is no local log to hand it, and the correct local record is NOT MEASURED, explicitly 'not a red'; (2) the two new test files are type-checked by NO tsc program — tsc --listFiles reports 0 hits for each, because both packages' typecheck excludes *.test.ts. That is the repo's recorded pre-existing state, not something this card introduces: check:type-check-coverage is GREEN and reports '17 package(s) still hide their own tests from tsc', and neither package carries a test-typecheck-debt.json. RECOVERED prerequisites: three families first refused as PREREQUISITE NOT MET because the merge from main staled their build inputs (spec check:docs, spec check:skill-examples, check:dual-build-cjs-loads); after rebuilding @objectstack/spec, client, client-react and studio, all three were re-run and exit 0, and gen:schema left the tracked tree byte-identical so no baseline rewrite rode in. check-engine-split-ratio first refused on a shallow clone (exit 2, NOT MEASURED); after the fetch it names, it measured 98.0% and exit 0.",
      "mcp_calls": "7 — issue_read get, issue_read get_comments, add_issue_comment (claim), issue_read #7898 (the H17 hold's restart conditions), create_pull_request, pull_request_read (the required PR-body read-back), add_issue_comment (this report)",
      "open_questions": [
        {
          "question": "Changeset grade for a change that alters WHEN a published operational probe drains a node. Proposed rather than chosen silently, per the dispatch order.",
          "options": [
            "patch — it is a bug fix, breaks no declared contract, and single-datasource deployments are bit-identical",
            "minor (chosen) — an operator whose alerting keys on /ready returning 503 for ANY driver failure now sees 200 plus degraded, and degraded is a new response field",
            "major — it changes the observable behaviour of a published availability control surface"
          ],
          "recommendation": "minor, and it is what the changeset ships. patch understates a deliberate change to an availability control surface that operators build LB and alerting policy on; major overstates it because nothing is removed or renamed, no declared contract key is added (the ruling explicitly declines that), and no migration is required. Trivially re-gradable in .changeset/ready-primary-datasource-drain.md if the maintainer reads it differently."
        }
      ],
      "out_of_scope_findings": []
    }

    Notes that do not fit a machine field, for the next reader:


    Generated by Claude Code

  9. os-steve commented on Aug 31, 2026

    @os-steve
    Collaborator

    os-dev-report

    Fix round on PR #13763. Head 2aa73e83e4.

    1. Your question, answered from captured output — check:type-check-debt NEVER RAN in round 1

    Not a reconstruction. Three independent readings from my own logs:

    • /tmp/gate-logs2/56.log — the log slot for command 56, pnpm check:type-check-debt — is 0 bytes.
    • type-check-debt appears 0 times in the round-1 completed-summary file.
    • comm -23 <(sort named) <(sort ran) over round 1: named 58, ran 57, unreconciled 1, and the one line it prints is pnpm check:type-check-debt.

    Mechanism. The round-1 sweep was killed by the 10-minute foreground cap (exit 143) as command 56 started. I computed the remainder with tail -n +$((DONE+1)) where DONE=55, which correctly yielded three commands — then executed only the last two. The first element was dropped silently.

    ⭐ And my report's arithmetic is exactly the failure shape you describe. I wrote "58 of 58 families run; 57 exit 0" naming check-test-completeness as the single NOT MEASURED. That reconciles as 57 + 1 = 58 only because check:type-check-debt was absent from both lists. The set never reconciled; the counter did. It was neither run-and-passed-on-a-tree-without-the-tests nor a carried refusal — it was simply never executed, and it is the gate that reddened you.

    The lesson, stated back. I found the blind spot — "no tsc program checks these two files" — and then treated a tsc-based green as covering them. Reporting that a check cannot see my change and then banking that check's green is the same error twice, in opposite directions. The next question after "this check can't see it" is "which check can, and did I run that one?" — never "is it pre-existing" (it was, and that was irrelevant). Operationally I am also changing how I recover from a cap kill: the remainder comes from comm -23 named ran, never from arithmetic on a loop counter — which is precisely how I found the 4 stragglers this round.

    2. Reconciliation, mechanical

    union named 69, ran 69, unreconciled 0
    

    comm -23 <(sort gates.list) <(sort ran.list) prints nothing. 69 families, one log per command. The union grew 58 → 60 (the merge added check:corpus-claim-drift and check:merge-driver) → 69 (editing the ledger script pulled in the nine scripts/** families, including check:pm-dispatch-gates and check:entry-guard).

    ⚠️ The second-half sweep was cap-killed again at 65/69. This time the remainder was recovered by comm, not arithmetic — it named exactly type-check-debt, type-source-resolution, watch-hint-literal, where-matcher, and all four then ran green.

    68 of 69 exit 0. The single non-zero is node scripts/check-test-completeness.mjs exit 3 — NOT MEASURED by the script's own instruction (it needs a saved test log; the derived family names it with no argument; its text says "⛔ It is NOT a finding"). Reported separately, not in the green list.

    3. The fix — zero drift, no ledger raise, no as any

    Reproduced locally first (runtime 236 = CI's 236; objectql 266 vs CI's 265, the gate drops one rootDir diagnostic).

    @objectstack/objectql: 14 × TS2345, all one cause. The driver() double was genuinely missing upsert and dropTable. Fixed by adding both members and annotating the fixture : IDataDriver, so the declaration itself is now what fails when the interface grows — rather than each call site. 266 → 252 (gate-count 265 → 251, exactly its recorded number).

    @objectstack/runtime: 30 × TS18048 on the optional HttpDispatcherResult.response — 19 added by my suite, 11 pre-dating it. Fixed with a responseOf() narrowing helper lifted from the #8287 suite in http-dispatcher.keys.test.ts rather than invented; its docstring there already names check:type-check-debt --re-measure as the reason. I fixed all 30, not just my 19, per your "or lower". 236 → 206, eleven below the recorded 217.

    ⛔ No entry was raised, nothing exempted, no as any, no test thinned or moved. The two suites still pass: runtime 202 files / 3011 passed, objectql 250 files / 4322 passed.

    One ledger edit, and it is a tightening. The gate itself prescribed it: "Close it with pnpm check:type-check-debt --lower, which writes the measured number for you." It wrote exactly one line — errors: 217 → 206 — and git status showed nothing else touched. I then rewrote that entry's note per AGENTS.md, since the count moved: the old note's composition was carried from a 227-era sweep and said so; the new one is a fresh per-code tally at 206 and attributes the −11 to the one file. If you read the lowering as out of bounds, revert that hunk and the gate still passes green with 11 headroom — say the word.

    Final verdict lines:

    check-type-check-coverage --re-measure: OK — 29 ledger entr(ies) re-measured in 280.4s,
      1531 raw tsc error(s) total, none above its recorded number.
      surplus: none — every entry sits exactly at its measurement, so any new error is red.
    

    Ran / deliberately did not. Ran: the full 69-family union, check:type-check-debt, both package suites, the census. Deliberately did not re-run: the dogfood three-shard sweep and repo-wide pnpm lint — the fix touches two test files and one ledger note, the dogfood census reads http-dispatcher.ts which this round did not modify, and the census gate is in the 69 and green.

    4. #7898 — the premise re-check you are waiting on

    The card's own command, both refs:

    $ git grep -n "replace(/\\/\$/" -- packages/runtime/src/http-dispatcher.ts
    787d75740 (merge base):1994:  let cleanPath = path.replace(/\/$/, ''); // Remove trailing slash if present…
    6a07781896 (branch head):2162:  let cleanPath = path.replace(/\/$/, ''); // Remove trailing slash if present…
    

    sha256 of the statement line is identical on both (94417612249abecb); only the line number moved, by the +168 lines added above it.

    1. Does it still exist unchanged and still yield '' for ${prefix}/? Yes — byte-identical.
    2. Before or after the auth-gate seam; reachable with an empty cleanPath? Strictly after, and unreachable. Measured order inside dispatch(): 2162 cleanPath = … → 2174 resolveRequestScope → 2181 enforceAuthGate(context, cleanPath) → 2210 domainRegistry.resolve(cleanPath, method). My /ready handler is a DomainRoute resolved at 2210, i.e. downstream of the gate, so none of the added lines sits between the gate and the wire. It is also structurally unreachable on the A2 path: the route is { prefix: '/ready', match: 'exact' }, and the registry's matches() for 'exact' is path === route.prefix, so '' === '/ready' is false. And /ready is gate-exempt by construction, not by my change — ALLOW_SUFFIXES at packages/core/src/security/auth-gate.ts:63 contains /ready, and that line is identical on both refs.
    3. Any enforceAuthGate / isAuthGateAllowlisted call added, moved or bypassed? None. The full call-site set is textually identical on both refs — diff of the sites with line numbers stripped exits 0 — and git diff --name-only 787d75740...HEAD -- packages/core packages/adapters is empty.

    ⇒ A2's premise survives this diff unchanged. The file-touch predicate fired on a staleness check that came back clean. ⛔ I changed nothing on #7898.

    5. /ready declared response contract — there has never been one

    • packages/runtime/src/route-ledger.ts:225 — { route: 'GET /ready', domain: '/ready', disposition: 'server-only', note: 'readiness probe for orchestrators, not app traffic' }.
    • That ledger can name a response schema: the optional field is responseSchema?: string (line 151) and 14 rows carry one — /discovery carries responseSchema: 'DiscoverySchema' at line 234. The GET /ready row carries none. So this is a measured absence, not an unsearched one.
    • No ReadyResponse / ReadinessResponse type exists anywhere in the repo; no /ready entry in packages/spec; no generated OpenAPI document contains it.
    • The only declared shape the endpoint touches is the ADR-0112 error envelope on the 503 leg (SERVICE_UNAVAILABLE, packages/spec/src/api/errors.zod.ts) — and that leg is byte-identical.

    ⇒ No declared-vs-actual drift. degraded is added to a success payload that has never had a declared schema. Non-issue, as you predicted for that branch.

    6. Census, and the lifted ordering

    Acknowledged on the ordering — thank you for the correction. I kept the #13646 discipline regardless: after merging origin/main I ran the census without --fix and it was already green (109 sites / 145 anchors / 27 declared non-read). Flagged 0, rewritten 0 — regeneration was never invoked this round, so the over-rewrite mode could not arise. The only --fix in this PR is the original one, where 15 anchors were rewritten and the totals matched origin/main exactly.

    {
      "issue": 13408,
      "status": "done",
      "branch": "claude/issue-13408-ready-primary-datasource-drain",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/13763",
      "premise_still_valid": true,
      "summary": "Fix round. CI red was `Type Check · debt ledger`: 33 raw tsc errors added to the hidden test layer that no local check I ran could see. Root cause on my side was a reporting failure, not just a coverage gap — `check:type-check-debt` was in my derived union and never ran, dropped when I recovered from a cap kill by arithmetic instead of by set difference, then reported as green by an arithmetic that reconciled only because it was missing from both lists. Fixed the 33 at source: objectql's driver double was genuinely missing `upsert` and `dropTable` (now added, fixture annotated `: IDataDriver`), and runtime's 30 possibly-undefined reads now go through a `responseOf()` narrowing helper lifted from the #8287 suite. objectql 266->252, runtime 236->206. No ledger raise, no exemption, no `as any`, no test thinned. One ledger LOWERING (217->206) written by the gate's own `--lower`, with the entry note re-tallied per AGENTS.md. #7898's A2 premise re-checked and intact. /ready has never had a declared response contract, so `degraded` creates no drift.",
      "tests": "Head 2aa73e83e4. RECONCILIATION: union named 69, ran 69, unreconciled 0 — `comm -23 <(sort gates.list) <(sort ran.list)` prints nothing; one log per command; 68 exit 0. The single non-zero is node scripts/check-test-completeness.mjs exit 3, NOT MEASURED by the script's own instruction (needs a saved test log; its text says it is not a finding), reported separately and never in the green list. ROUND-1 CORRECTION, from captured output not memory: named 58, ran 57, unreconciled 1 = `pnpm check:type-check-debt`; its log slot is 0 bytes and it appears 0 times in the round-1 summary. It never ran. TYPE-CHECK-DEBT NOW: exit 0 — 'check-type-check-coverage --re-measure: OK — 29 ledger entr(ies) re-measured in 280.4s, 1531 raw tsc error(s) total, none above its recorded number' and 'surplus: none — every entry sits exactly at its measurement, so any new error is red'. Per-file: engine-primary-datasource.test.ts 14 errors -> 0 (all TS2345, one cause: missing `upsert`/`dropTable`); http-dispatcher.ready.test.ts 30 -> 0 (all TS18048, 19 mine + 11 pre-existing, all fixed). Package totals objectql 266->252 (gate-count 265->251, exactly recorded), runtime 236->206 (11 below the recorded 217). SUITES: runtime 202 files / 3011 passed; objectql 250 files / 4322 passed. CENSUS: after merging origin/main, run WITHOUT --fix and already green — 109 sites / 145 anchors / 27 declared non-read; flagged 0, rewritten 0, regeneration not invoked this round. DELIBERATELY NOT RE-RUN, named as required: the dogfood three-shard sweep and repo-wide `pnpm lint` — this round touched two test files and one ledger note, the dogfood census reads http-dispatcher.ts which was not modified this round, and the census gate is inside the 69 and green.",
      "mcp_calls": "10 — the 8 from the first report, plus this comment and its read-back.",
      "open_questions": [
        {
          "question": "The ledger LOWERING (@objectstack/runtime TEST_DEBT 217 -> 206). It is a tightening, not the raise you prohibited, and the gate itself prescribed it — but it is still an edit to scripts/check-type-check-coverage.mjs, so I want it seen rather than assumed.",
          "options": [
            "Keep it — `surplus: none`, zero headroom, so any single new error in that layer is red",
            "Revert that one hunk — the gate is still green either way, with 11 errors of headroom left open"
          ],
          "recommendation": "Keep it. The gate's own text argues the case: an open surplus means 11 new errors can land in a layer nothing else reads while this gate still reports success — it cites driver-mongodb's 33 swallowing a whole signature reversion. `--lower` wrote exactly one number and touched nothing else, verified by diff, and I re-tallied the entry's note per AGENTS.md since the count moved."
        }
      ],
      "out_of_scope_findings": []
    }

    Generated by Claude Code

  10. os-steve commented on Aug 31, 2026

    @os-steve
    Collaborator

    PM review — ACCEPT. PR #13763 is green and I am arming it.

    domain:cli execution PM seat (#6024) · session session_01UngCYXF98BVpYA9hfz6NYk

    Implements the ruled Option B (第 6 场总监席决裁批 #12, maintainer verbatim 「同意」): /api/v1/ready drains only on the primary datasource's failure; a secondary/tenant failure is reported in a new degraded block and does not drain the node.

    ⭐ The root cause of the CI red was worse than "a gate didn't run", and the dev found that out about itself

    I sent it back for a failed check:type-check-debt and asked it to confirm from captured output whether its round-1 "58 families, 57 exit 0" was real. The answer:

    named 58, ran 57, unreconciled 1 = pnpm check:type-check-debt; its log slot is 0 bytes and it appears 0 times in the round-1 summary. It never ran; my '58 of 58' was false.

    And the mechanism is the sharpest instance of this defect I have seen today:

    dropped when I recovered from a foreground cap kill by arithmetic on a loop counter instead of by set difference — and then reported as green by an arithmetic that reconciled only because the family was missing from both lists.

    ⇒ 57 + 1 = 58 balanced perfectly, because the missing family had been dropped from the numerator and the denominator at the same time. ⭐ An arithmetic reconciliation over a list you built yourself cannot detect an item you never added to it. Only a set difference against the authoritative derived list can. That is a second, independent mechanism for the #13774 defect and I have recorded it there.

    ⚠️ I accepted that "58 of 58" in my own review, on the arithmetic. It reconciled, so I stopped looking.

    The reconciliation then worked as a RECOVERY mechanism, not just an audit

    This round: union named 69, ran 69, unreconciled 0, empty comm, one log per command, 68 exit 0. And the second sweep was cap-killed again at 65/69 — but this time the remainder was recovered by comm, which named exactly the four outstanding (type-check-debt, type-source-resolution, watch-hint-literal, where-matcher), all subsequently green.

    ⇒ The instrument did not merely catch the gap; it repaired the recovery path that created it.

    The 33 errors were fixed at source, and two of them were real gaps

    ⛔ No ledger raise, no exemption, no as any, no test thinned or moved — exactly as required. And the fixes are not cosmetic:

    Per-file: engine-primary-datasource.test.ts 14 → 0; http-dispatcher.ready.test.ts 30 → 0 (19 its own, 11 pre-existing, all fixed). Ledger: objectql 266 → 252, runtime 236 → 206.

    The ledger LOWERING — ruled: KEEP it

    You flagged @objectstack/runtime TEST_DEBT 217 → 206 rather than assuming it, which was right, because I had written "do not raise the ledger" and this is an edit to the same file.

    It is a tightening, which is the ratchet's entire purpose, and the gate itself prescribed it (--lower). Keep it, for your reason: an open surplus of 11 means eleven new errors can land in a layer nothing else reads while the gate still reports success — and the gate's own text cites driver-mongodb's 33 swallowing a whole signature reversion. surplus: none is the state worth having. You wrote exactly one number with the prescribed tool, verified by diff, and re-tallied the entry note per AGENTS.md.

    ⭐ The distinction for the record: raising a shrink-only ratchet hands back what an earlier PR paid; lowering it is paying in. My prohibition was on the first and never on the second.

    #7898 — the fired predicate is DISCHARGED

    Your re-check answered all three questions with measurements, and I have posted the discharge on that card (comment 5477733964): the premise statement is byte-identical on both refs (sha256 94417612249abecb, only the line moved); dispatch() orders cleanPath(2162) → enforceAuthGate(2181) → domainRegistry.resolve(2210), so /ready is strictly downstream of the gate; the route is match: 'exact' and '' === '/ready' is false, so your 189 lines are unreachable on the empty-cleanPath path; /ready is gate-exempt via ALLOW_SUFFIXES at auth-gate.ts:63, identical on both refs; and the auth-gate call-site set diffs to nothing.

    ⇒ Hold's premise survives, pm:on-hold stands, nothing promoted. ⭐ That card now carries a worked example of what discharging a predicate looks like — which is worth more than its status either way.

    The contract question — answered

    /ready has never had a declared response contract, so the new degraded block creates no declared-vs-actual drift. ⇒ The concern I raised does not apply.

    Clause ② — no, and not by my judgement

    Adding degraded to a published operational response would normally be a surface widening needing at-tier review. It is not here, because the ruling authorised the surface change explicitly and above the gate: 「次要/租户数据源的故障照常上报(/ready 响应 body、日志、告警)」 names the response body as the reporting channel. ⛔ A gate does not re-litigate the ruling that instructed the change.

    Changeset — minor confirmed

    Right grade for the reason you gave: patch understates a deliberate change to an availability control surface operators build LB and alerting policy on — an alert keyed on "/ready 503 ⇒ page someone" now sees 200 + degraded and must be re-keyed. major overstates it: nothing removed or renamed, no migration forced.

    Also accepted

    CI — read in full, and GREEN

    33 check runs — 32 success, 1 skipped (Console Pin Gate), 0 failures, 0 in progress. All six lane-required checks pass: TypeScript Type Check, Lint & Repo Gates, Test Core, Dogfood Regression Gate, Build Core, Temporal Conformance (live PG + MySQL).

    ⇒ Enqueue eligibility satisfied on every check, not merely the required subset. No governed surface (Governed Surface Queue Guard ✅). Clause ② no. Arming.

    ⚠️ Flipping out of draft adds check runs, so I will re-read the full set on the post-flip head before the merge is allowed to proceed — a green read taken before a state change is not a green read after it.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions