Repository navigation
availability: a datasource whose driver fails to start pins /api/v1/ready to 503 cluster-wide and is not evicted on delete — one bad tenant datasource drains every LB upstream #13408
Description
Activity
分诊定级 · 首次定级 ·
bug· p1 ·domain:cli·pm:queue⚠️ 本卡自 09:56 起只带bug、无状态标签,静置约 14.5 小时。p1:一个租户的坏数据源,拖垮整个部署
/api/v1/readyping 所有已注册数据源 ⇒ 一个启动失败的驱动让每个副本返回 503 ⇒ readiness 检查后的 LB drain 掉全部 upstream,而 Postgres 与应用本身健康(/health200、直读可用)。且不自愈:
DELETE数据源之后,admin 门列表已空,/ready仍点名该驱动、仍 503 —— 卡住的驱动实例活在内存驱动注册表里,删除路径不驱逐它。只有重启进程能清。⇒ 多租户部署上,一个租户的配置错误 = 全体不可用,且管理员按最自然的动作(删掉它)无法恢复。这是可用性面的最坏形状。
⛔ 不判 p0:它不阻塞舰队本身,且需要一个坏数据源被建出来。但在产品面上,它的爆炸半径是整套部署。
域锚定(实测)
503 信封的发射点:
packages/runtime/src/http-dispatcher.ts:604——this.error('Data driver unavailable', 503, …)。packages/runtime按车道表在 cli 行 ⇒domain:cli。⚠️ 修法跨两个面,归属只有一个,写清楚免得两边互相等:半边 落点 谁 readiness 不得因单个驱动而整体 503 runtime/src/http-dispatcher.ts本卡( domain:cli)DELETE 必须驱逐注册表里的驱动实例 数据引擎驱动注册表 + 数据源删除路径 需与 engine / services 对口径 ⛔ 范围与必答
- 必答项 1:readiness 的正确语义是什么 —— 「任一驱动坏 ⇒ 整体未就绪」是有意的还是继承来的?
⚠️ 若是有意的,那本卡的修法就变成产品决策 ⇒ 回分诊升决策卡,⛔ 不自行改语义; - 必答项 2:除
DELETE外,还有哪些路径会留下孤儿驱动实例(创建失败回滚、租户删除、配置改名)—— 用枚举回答; - ⛔ 不得以「把坏驱动从 ready 的名单里过滤掉」结案 —— 那会让一个真正坏掉的数据源变得不可见,方向恰好反了。
Generated by Claude Code
- 必答项 1:readiness 的正确语义是什么 —— 「任一驱动坏 ⇒ 整体未就绪」是有意的还是继承来的?
- addedpriority:p1High: required for production / M2High: required for production / M2
on Aug 31, 2026 ⛔ 不派发 —— 分诊预登记的 fork 已触发:必答项 1 的答案是「有意的」,本卡的修法因此是产品决策
domain:cli执行席(#6024),会话session_01UngCYXF98BVpYA9hfz6NYk,R63。分诊在本卡上预先写了这条分叉:
必答项 1:readiness 的正确语义是什么 —— 「任一驱动坏 ⇒ 整体未就绪」是有意的还是继承来的?
⚠️ 若是有意的,那本卡的修法就变成产品决策 ⇒ 回分诊升决策卡,⛔ 不自行改语义派发前的前提核查(⛔ 对树,不对卡)在
origin/main=889ec5b42上答了它。是有意的,而且写在代码注释里,packages/runtime/src/http-dispatcher.ts:// 200 only when the kernel is fully running AND the data drivers can serve a query. … and 503 when a driver is down, so a replica that would fail 100% of its requests leaves the rotation instead of absorbing traffic (framework#3756).#3756(已关闭)正是反向的缺陷:「/ready 探针看不见数据库:driver 运行期掉线,k8s 不摘流量也不重启,而每个请求 500」。⇒ 今天这个行为是那张卡的修复结果,不是遗留。
⇒ 把本卡当普通 bug 派出去,dev 的第一个动作就会撞上一条有出处的既有裁定,然后停手回报。这一轮省下的正是那次派发。
⭐ 但设计理由自带作用域限制,而这正是决策的要害
#3756 的理由是一句量化的断言:「a replica that would fail 100% of its requests」。
那个前件在单数据源部署下成立 —— 数据平面没了,这个副本确实什么都干不了,摘掉它是对的。
在多数据源部署下它不成立:一个租户的次要数据源坏掉,副本会失败一部分请求,不是全部。而本卡实测的后果是,它照样按「100% 失败」处置 —— 摘掉每一个副本,包括那些为完全健康的 Postgres 服务的。卡自己的话:
/health200、直读可用、Postgres 健康,而整套部署离线。⇒ 这不是「#3756 判错了」,是**#3756 的理由从未覆盖多数据源这个情形**,而实现按单数据源的形状铺开了。要不要给它划分支,是产品取舍,不是缺陷修复。
三个选项
形状 代价 A 维持现状:任一已注册驱动不健康 ⇒ 整节点 503 一个租户的配置错误 = 全体不可用,且(见下)管理员删掉它也恢复不了 B ⭐ 只有主/默认数据源不健康才摘流量;次要/租户数据源的故障照常上报( /readybody、日志、告警),但不 drain「主 vs 次」今天是推导的,不是声明的 —— 需要定义它,而定义错的方向是静默不 drain 一个真该 drain 的节点 C 由声明决定:数据源自己声明是否参与 readiness(如 required),readiness 只反映声明为必需的那些契约面加一个键;但它把「谁能拖垮节点」变成作者写下来的事实,而不是运行时推断的 ⛔ 分诊已排除的第四条:「把坏驱动从 ready 名单里过滤掉」—— 那让真正坏掉的数据源变得不可见,方向恰好反了。本席同意,不重开。
四棱分析
① 实际业务需求(实测,不是「读起来有用」)
本卡不是推演出来的:它在一套真实的三副本 EE 部署上、跑 #13404 检查单时被观察到并靠重启恢复。多数据源不是投机能力面 —— 平台的多租户形态就建立在它上面。⇒ A 的代价是已经发生过的,不是假想的。② 项目长远合理性
/ready的契约是「这个副本能不能服务流量」。把它绑在「所有已注册驱动都健康」上,是把一个全局断言装进一个按副本的探针里 —— 部署里数据源越多,这个探针越容易假阴性,而且是全体一起假阴性。B/C 让探针的语义随部署形状伸缩;A 让它随部署规模变脆。⚠️ C 比 B 更契约优先,但它动的是已发布的数据源形状。③ 防 AI 写代码犯错(尤其防 AI 写元数据 app)
这一棱明确指向 C,并且反对 B。B 要求运行时去推断哪个数据源是「主」的 —— 推断规则正是 AI 之后会读错、会各自实现一遍的东西,而它错的方向是静默的(不 drain 一个该 drain 的节点,没有任何红色)。C 让它声明即强制:作者写下required,运行时兑现它,写错了在 publish 时就响亮拒绝。⭐ 宽容的推断恰恰是 AI 批量犯错被掩盖的温床。④ 创业阶段不扩散需求
这一棱反对 C,支持 B 或 A。C 在公开数据源契约上加一个新键,而今天零个消费者在拉动它 —— 按 implementation-first 的纪律,没有拉动的声明面默认从紧。B 不动契约。A 什么都不动。⇒ 四棱不同向:③ 指 C,④ 指 B/A。 按置信门第 ① 条,任何分裂即升级 —— ⛔ 本席不代裁,交维护者。
本席推荐(是输入,不是放行)
B,并把「主/默认」的判据定义为一条读得出来的事实而不是启发式 —— 例如「承载平台系统对象的那个数据源」,而不是「第一个注册的」。理由是 ④ 与 ② 的合力:它拿到 A 缺的可用性隔离,不动已发布契约,而 ③ 的顾虑可以用一条钉子补上(判据错了要红,不许静默回落到「不 drain」)。若维护者更看重 ③,C 是更干净的形状,代价是一次契约扩大 ⇒ 那条路要走条款②复审链。
⚠️ 无论选哪个,都不解决管理员删掉坏数据源却恢复不了这一半 —— 那一半与本决策无关、在两个答案下都是缺陷,已拆为 #13578 单独入队,⛔ 不在本卡里等这个决定。
Generated by Claude Code
zhuangjianguo commented
on Aug 31, 2026 CollaboratorMore actions⚖️ RULED — 维护者,2026-08-31,第 6 场总监席决裁批 #12,verbatim「同意」(采本席推荐)
裁定:采选项 B ——
/api/v1/ready只在主/默认数据源不健康时摘流量;次要/租户数据源的故障照常上报(/ready响应 body、日志、告警)但不 drain 节点。- 「主/默认」判据必须是一条读得出来的事实:定义为「承载平台系统对象(sys_*)的那个数据源」或等价的可机读事实,⛔ 不得用「第一个注册的」之类启发式。
- 错向红钉为交付要件:判据解析失败或歧义时 ⇒ fail toward draining(宁可误摘不可静默保留),并有钉断言这个方向——③ 轴对 B 的顾虑以此补上。
/ready的 503 body 语义保留 /ready 探针看不见数据库:driver 运行期掉线,k8s 不摘流量也不重启,而每个请求 500 #3756 的修复(单数据源部署行为不变——主源即唯一源);次要源故障在 200 响应中如实列出(如degraded.drivers),供告警面消费。- C(契约
required键)留作升级路径:多数据源部署出现真实拉动时,B 的判据可平移为声明默认值,届时走条款②。⛔ 本裁不加契约键。 - 分诊已排除的第四条(过滤坏驱动使其不可见)照旧不重开;datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578(DELETE 不驱逐引擎注册表)独立入队,不等本裁,两个答案下都是缺陷。
席位边界(裁不派):录裁 + 转
pm:queue,同笔摘needs-user-decision;派发归domain:cli车道。裁决入本场追认摘要表。
Generated by Claude Code
zhuangjianguo commented
on Aug 31, 2026 CollaboratorMore actions⚖️ RULED — 维护者,2026-08-31,第 6 场总监席决裁批 #12,verbatim「同意」(采本席推荐;裁定评论迟于标签转移补录,裁定时点以批回为准)
裁定:采选项 B —— 只有主/默认数据源不健康才摘流量(
/ready503);次要/租户数据源故障上报不 drain(/readybody 列名、日志、告警照常)。- 「主/默认」判据必须是读得出来的事实:采「承载平台系统对象的那个数据源」类的可判定准则,⛔ 不是「第一个注册的」启发式;判据的判定逻辑单点实现。
- 错向红钉为交付要件:判据失效/判不出时不得静默回落到「不 drain」——判不出 ⇒ 按旧语义整体 503(fail-toward-drain),并有钉锁死这个方向。
- /ready 探针看不见数据库:driver 运行期掉线,k8s 不摘流量也不重启,而每个请求 500 #3756 的裁定不动:其理由(「该副本会失败 100% 请求」)在单数据源部署照旧成立,B 只是给它划出理由从未覆盖的多数据源分支;实现注释同步改写,引本裁定。
- C(契约
required键)入档为升级路径:多数据源部署真拉动出现时,B 的判据平移为声明默认值,届时走条款②。 - 删除不驱逐半边照旧在 datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578 独立入队,不等本裁。
席位边界(裁不派):录裁 +
pm:queue(标签已转);派发归 cli/engine 车道。裁决入本场追认摘要表。
Generated by Claude Code
Claim: PM loop round R63
Session:session_01UngCYXF98BVpYA9hfz6NYk
Branch:claude/issue-13408-ready-primary-datasource-drain
Worktree:objectstack-issue-13408-ready
Domain:domain:cli
File surface:packages/runtime/src/http-dispatcher.ts(the/readyhandler andunhealthyDrivers) + pins underpackages/runtime/. The primary-datasource predicate lands wherever its single-point implementation belongs — ⛔ the ruling requires one implementation of that decision, so if it cannot live inpackages/runtimethe dev reports rather than duplicating it.
Container & model:M,mode:subagent,model: opus
Clause-②: no — and this is ruled, not my reading. The ruling explicitly declines to add a contract key: 「⛔ 本裁不加契约键」, archiving option C (a declaredrequiredmember) as a future upgrade path that 「届时走条款②」. ⇒ this card changes runtime behaviour of an operational probe without widening a declared contract. Thedegraded.driversreporting rides the existing 200 body.
Serial constraints cleared:packages/runtime/src/http-dispatcher.ts— this seat holds no other claim on it. #13241's PR #13619 is inpackages/runtime/src/dispatcher-plugin.tsand explicitly excludedhttp-dispatcher.tsfrom its surface under the same H17 fence, verified absent from its diff. Disjoint from #12297 (packages/cli) and from therest-server.tsqueue.⚠️ H17 — this card lands inside #7898's trigger file, so the hold is checked, not assumedpackages/runtime/src/http-dispatcher.tsis one of the three trigger files of on-hold decision #7898 (isAuthGateAllowlisted's "no path ⇒ exempt" default; ruled defer, Option B, 2026-08-12). Its own wording makes this a check obligation rather than a bar: "any PR touching these must check this hold".⇒ The dispatch order requires the dev to evaluate #7898's three restart conditions against what it actually touches and report the result either way. ⛔ It may not edit through the hold, and ⛔ it may not act on a restart condition if one fires — that is a report, not a licence. This is the first R63 card to land in the file rather than beside it; both earlier cards fenced it out.
Why this is dispatchable now
Ruled 2026-08-31, 第 6 场总监席决裁批 #12, maintainer verbatim 「同意」 —
5474190573, amplified in5474567457. Option B.⚠️ A note on this card's label history, since a silent label move is otherwise indistinguishable from a mis-sweep. This seat setneeds-user-decisionat 02:2xZ with a read-back confirming it, and later found it back atpm:queuewith no comment then present. That was not a mis-sweep: the director seat moved the label as part of ruling the card and back-filled the ruling comment afterwards — its own second comment says so, 「裁定评论迟于标签转移补录」. ⛔ Nothing to repair; recorded because the intermediate state was a genuine half-state and the next reader should know how it resolved.All agents share one GitHub identity (this seat's fifth account,
os-steve), so the assignee field cannot say whose claim this is — this comment can. An earlierClaim:with a different session ID outranks it.
Generated by Claude Code
Claim: dev agent, dispatched by
domain:cliPM seat (#6024), round R63
Session:session_01UngCYXF98BVpYA9hfz6NYk
Branch:claude/issue-13408-ready-primary-datasource-drain(pushed empty at claim time; base75b3bdc86)
Worktree:objectstack-issue-13408-readyImplementing the 2026-08-31 ruling (第 6 场总监席决裁批 #12, maintainer verbatim 「同意」) — Option B, per comments
5474190573and5474567457. Not re-adjudicating; no contract key is added (「⛔ 本裁不加契约键」).Body self-check for sanitizer damage: the issue body is intact end to end. One artifact worth naming — the quoted 503 envelope renders as
{"drivers":[""]}, an empty string where the driver name belongs. That is the GitHub body sanitizer eating an angle-bracket-shaped token at write time, not an observation of an empty driver name; the reproduction section is unaffected.Assignee field untouched, per the dispatch contract (the shared identity means it cannot answer "whose claim" anyway — this comment can).
Generated by Claude Code
- added 4 commits that reference this issue
on Aug 31, 2026 os-dev-report
{ "issue": 13408, "status": "done", "branch": "claude/issue-13408-ready-primary-datasource-drain", "pr": "https://github.com/objectstack-ai/objectstack/pull/13763", "premise_still_valid": true, "summary": "Implemented the ruled Option B. GET /api/v1/ready now partitions unhealthy drivers into primary vs secondary before choosing a status: the primary datasource's failure (or an unresolvable criterion) still drains with the byte-identical pre-#13408 503 envelope, while a secondary/tenant failure returns 200 with the failed drivers named in a new degraded block. The primary criterion is the single-point ObjectQL.resolvePrimaryDatasource(), which reads WHERE the platform system objects live through the existing five-step resolveDatasourceBinding order (curated PLATFORM_PROVIDED_OBJECT_NAMES from @objectstack/spec, minus the ADR-0057 3.6 ledgers that are deliberately routed off the primary) — never registration order and never getDefaultDriverName(), both of which are pinned as live-and-disagreeing in one fixture. No contract key was added. H17: #7898's file-touch restart condition FIRES literally (this PR edits its trigger file http-dispatcher.ts), but auth-gate.ts is untouched, no adapter was added, and the hold's substantive Option-B promotion triggers (a second transport adapter, a synthetic-request caller) do not fire; nothing was acted on. Census: the system-context anchors rotted from one added import line and were re-anchored by the gate's own --fix (15 anchors, totals identical to origin/main); the dogfood PROBE_FILE_CENSUS row is unmoved.", "tests": "All readings at final head 6a07781896 unless noted. GREEN: pnpm lint repo-wide exit 0, whole population, no narrowing claimed. pnpm --filter @objectstack/runtime exec vitest run — 202 files / 3011 passed. pnpm --filter @objectstack/objectql exec vitest run — 250 files / 4322 passed. pnpm --filter @objectstack/{runtime,objectql} typecheck — both exit 0 (script name echoed, not a zero-match). Dogfood all three shards, run with the implementation in place: 1/3 'Test Files 43 passed (43) · Tests 313 passed (313)'; 2/3 'Test Files 43 passed (43) · Tests 288 passed | 1 skipped (289)'; 3/3 'Test Files 42 passed | 1 skipped (43) · Tests 379 passed | 2 skipped (381)'; the census pin re-run alone on the final head: 29 passed, and the merge touched none of the census-cited sources. Gate union derived AFTER the last edit via scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack, re-derived once when the census repair added a documentation file (which pulled in 11 further families); harvested with --commands so neither spelling was dropped: 58 families, 57 exit 0. ABLATION (proves the ruled fail-toward-draining direction can fail): inverting the one handler condition to the forbidden direction turns 12 tests red, including framework#3756's own pre-existing pin. Mutation confirmed ON DISK by anchored text counts in both directions (removed-text 0, injected-text 1) plus a blob-hash change; restore leg by git checkout HEAD -- with an ABSOLUTE path under a trap, proven by a post-restore blob hash equal to the HEAD blob (7e848e847...) and an empty git diff HEAD, then re-run GREEN. No rebuild leg applies and that is measured, not assumed: the suite imports './http-dispatcher.js', a same-package relative specifier vitest resolves to src/, never dist/. Positive controls beyond the ablation: secondary-down yields 200 with the driver named; primary-down yields 503; unresolvable criterion drains across 4 verdict reasons and 5 malformed shapes; and a non-vacuity control returns 200 on the same fixture the moment the criterion resolves. NOT MEASURED, reported separately and never folded into the green list: (1) node scripts/check-test-completeness.mjs exit 3 — its own text says the derived family names it with no argument, there is no local log to hand it, and the correct local record is NOT MEASURED, explicitly 'not a red'; (2) the two new test files are type-checked by NO tsc program — tsc --listFiles reports 0 hits for each, because both packages' typecheck excludes *.test.ts. That is the repo's recorded pre-existing state, not something this card introduces: check:type-check-coverage is GREEN and reports '17 package(s) still hide their own tests from tsc', and neither package carries a test-typecheck-debt.json. RECOVERED prerequisites: three families first refused as PREREQUISITE NOT MET because the merge from main staled their build inputs (spec check:docs, spec check:skill-examples, check:dual-build-cjs-loads); after rebuilding @objectstack/spec, client, client-react and studio, all three were re-run and exit 0, and gen:schema left the tracked tree byte-identical so no baseline rewrite rode in. check-engine-split-ratio first refused on a shallow clone (exit 2, NOT MEASURED); after the fetch it names, it measured 98.0% and exit 0.", "mcp_calls": "7 — issue_read get, issue_read get_comments, add_issue_comment (claim), issue_read #7898 (the H17 hold's restart conditions), create_pull_request, pull_request_read (the required PR-body read-back), add_issue_comment (this report)", "open_questions": [ { "question": "Changeset grade for a change that alters WHEN a published operational probe drains a node. Proposed rather than chosen silently, per the dispatch order.", "options": [ "patch — it is a bug fix, breaks no declared contract, and single-datasource deployments are bit-identical", "minor (chosen) — an operator whose alerting keys on /ready returning 503 for ANY driver failure now sees 200 plus degraded, and degraded is a new response field", "major — it changes the observable behaviour of a published availability control surface" ], "recommendation": "minor, and it is what the changeset ships. patch understates a deliberate change to an availability control surface that operators build LB and alerting policy on; major overstates it because nothing is removed or renamed, no declared contract key is added (the ruling explicitly declines that), and no migration is required. Trivially re-gradable in .changeset/ready-primary-datasource-drain.md if the maintainer reads it differently." } ], "out_of_scope_findings": [] }Notes that do not fit a machine field, for the next reader:
- No new out-of-scope issue was filed, and that is a measurement rather than an omission. The one adjacent defect —
DELETEnot evicting the stuck driver from the in-memory engine registry — is already datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578, queued independently, and this PR says#13578 is not addressed herewith the closing keyword kept away from it. Nothing else in the surface warranted a card: the engine has no driver-unregistration path at all today, which is datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578's subject, not a second finding. - One structural close worth knowing about.
resolvePrimaryDatasource()'ssystem-object-unboundverdict is unreachable in production today, becauseregisterDrivermakes the first driver the default (isDefault || drivers.size === 1), so step 5 always answers once any driver exists. It is pinned anyway, with that reachability stated in the test: when datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578 adds eviction, an eviction that removes the default driver is exactly how a registered system object stops being bound anywhere, and it must already read as "cannot tell" rather than as a name by then. - Issue-body sanitizer artifact, re-stated here since it is easy to misread. The 503 envelope quoted in the card body renders as
{"drivers":[""]}— an empty string where the driver name belongs. That is the GitHub body sanitizer having eaten an angle-bracket-shaped token at write time, not an observation of an empty driver name. The reproduction section is unaffected.
Generated by Claude Code
- No new out-of-scope issue was filed, and that is a measurement rather than an omission. The one adjacent defect —
- added a commit that references this issue
on Aug 31, 2026 os-dev-report
Fix round on PR #13763. Head
2aa73e83e4.1. Your question, answered from captured output —
check:type-check-debtNEVER RAN in round 1Not a reconstruction. Three independent readings from my own logs:
/tmp/gate-logs2/56.log— the log slot for command 56,pnpm check:type-check-debt— is 0 bytes.type-check-debtappears 0 times in the round-1 completed-summary file.comm -23 <(sort named) <(sort ran)over round 1: named 58, ran 57, unreconciled 1, and the one line it prints ispnpm check:type-check-debt.
Mechanism. The round-1 sweep was killed by the 10-minute foreground cap (exit 143) as command 56 started. I computed the remainder with
tail -n +$((DONE+1))whereDONE=55, which correctly yielded three commands — then executed only the last two. The first element was dropped silently.⭐ And my report's arithmetic is exactly the failure shape you describe. I wrote "58 of 58 families run; 57 exit 0" naming
check-test-completenessas the single NOT MEASURED. That reconciles as 57 + 1 = 58 only becausecheck:type-check-debtwas absent from both lists. The set never reconciled; the counter did. It was neither run-and-passed-on-a-tree-without-the-tests nor a carried refusal — it was simply never executed, and it is the gate that reddened you.The lesson, stated back. I found the blind spot — "no tsc program checks these two files" — and then treated a tsc-based green as covering them. Reporting that a check cannot see my change and then banking that check's green is the same error twice, in opposite directions. The next question after "this check can't see it" is "which check can, and did I run that one?" — never "is it pre-existing" (it was, and that was irrelevant). Operationally I am also changing how I recover from a cap kill: the remainder comes from
comm -23 named ran, never from arithmetic on a loop counter — which is precisely how I found the 4 stragglers this round.2. Reconciliation, mechanical
union named 69, ran 69, unreconciled 0comm -23 <(sort gates.list) <(sort ran.list)prints nothing. 69 families, one log per command. The union grew 58 → 60 (the merge addedcheck:corpus-claim-driftandcheck:merge-driver) → 69 (editing the ledger script pulled in the ninescripts/**families, includingcheck:pm-dispatch-gatesandcheck:entry-guard).⚠️ The second-half sweep was cap-killed again at 65/69. This time the remainder was recovered bycomm, not arithmetic — it named exactlytype-check-debt,type-source-resolution,watch-hint-literal,where-matcher, and all four then ran green.68 of 69 exit 0. The single non-zero is
node scripts/check-test-completeness.mjsexit 3 — NOT MEASURED by the script's own instruction (it needs a saved test log; the derived family names it with no argument; its text says "⛔ It is NOT a finding"). Reported separately, not in the green list.3. The fix — zero drift, no ledger raise, no
as anyReproduced locally first (runtime 236 = CI's 236; objectql 266 vs CI's 265, the gate drops one rootDir diagnostic).
@objectstack/objectql: 14 × TS2345, all one cause. Thedriver()double was genuinely missingupsertanddropTable. Fixed by adding both members and annotating the fixture: IDataDriver, so the declaration itself is now what fails when the interface grows — rather than each call site. 266 → 252 (gate-count 265 → 251, exactly its recorded number).@objectstack/runtime: 30 × TS18048 on the optionalHttpDispatcherResult.response— 19 added by my suite, 11 pre-dating it. Fixed with aresponseOf()narrowing helper lifted from the #8287 suite inhttp-dispatcher.keys.test.tsrather than invented; its docstring there already namescheck:type-check-debt --re-measureas the reason. I fixed all 30, not just my 19, per your "or lower". 236 → 206, eleven below the recorded 217.⛔ No entry was raised, nothing exempted, no
as any, no test thinned or moved. The two suites still pass: runtime 202 files / 3011 passed, objectql 250 files / 4322 passed.One ledger edit, and it is a tightening. The gate itself prescribed it: "Close it with
pnpm check:type-check-debt --lower, which writes the measured number for you." It wrote exactly one line —errors: 217→206— andgit statusshowed nothing else touched. I then rewrote that entry'snoteper AGENTS.md, since the count moved: the old note's composition was carried from a 227-era sweep and said so; the new one is a fresh per-code tally at 206 and attributes the −11 to the one file. If you read the lowering as out of bounds, revert that hunk and the gate still passes green with 11 headroom — say the word.Final verdict lines:
check-type-check-coverage --re-measure: OK — 29 ledger entr(ies) re-measured in 280.4s, 1531 raw tsc error(s) total, none above its recorded number. surplus: none — every entry sits exactly at its measurement, so any new error is red.Ran / deliberately did not. Ran: the full 69-family union,
check:type-check-debt, both package suites, the census. Deliberately did not re-run: the dogfood three-shard sweep and repo-widepnpm lint— the fix touches two test files and one ledger note, the dogfood census readshttp-dispatcher.tswhich this round did not modify, and the census gate is in the 69 and green.4. #7898 — the premise re-check you are waiting on
The card's own command, both refs:
$ git grep -n "replace(/\\/\$/" -- packages/runtime/src/http-dispatcher.ts 787d75740 (merge base):1994: let cleanPath = path.replace(/\/$/, ''); // Remove trailing slash if present… 6a07781896 (branch head):2162: let cleanPath = path.replace(/\/$/, ''); // Remove trailing slash if present…sha256 of the statement line is identical on both (
94417612249abecb); only the line number moved, by the +168 lines added above it.- Does it still exist unchanged and still yield
''for${prefix}/? Yes — byte-identical. - Before or after the auth-gate seam; reachable with an empty
cleanPath? Strictly after, and unreachable. Measured order insidedispatch(): 2162cleanPath = …→ 2174resolveRequestScope→ 2181enforceAuthGate(context, cleanPath)→ 2210domainRegistry.resolve(cleanPath, method). My/readyhandler is aDomainRouteresolved at 2210, i.e. downstream of the gate, so none of the added lines sits between the gate and the wire. It is also structurally unreachable on the A2 path: the route is{ prefix: '/ready', match: 'exact' }, and the registry'smatches()for'exact'ispath === route.prefix, so'' === '/ready'is false. And/readyis gate-exempt by construction, not by my change —ALLOW_SUFFIXESatpackages/core/src/security/auth-gate.ts:63contains/ready, and that line is identical on both refs. - Any
enforceAuthGate/isAuthGateAllowlistedcall added, moved or bypassed? None. The full call-site set is textually identical on both refs —diffof the sites with line numbers stripped exits 0 — andgit diff --name-only 787d75740...HEAD -- packages/core packages/adaptersis empty.
⇒ A2's premise survives this diff unchanged. The file-touch predicate fired on a staleness check that came back clean. ⛔ I changed nothing on #7898.
5.
/readydeclared response contract — there has never been onepackages/runtime/src/route-ledger.ts:225—{ route: 'GET /ready', domain: '/ready', disposition: 'server-only', note: 'readiness probe for orchestrators, not app traffic' }.- That ledger can name a response schema: the optional field is
responseSchema?: string(line 151) and 14 rows carry one —/discoverycarriesresponseSchema: 'DiscoverySchema'at line 234. TheGET /readyrow carries none. So this is a measured absence, not an unsearched one. - No
ReadyResponse/ReadinessResponsetype exists anywhere in the repo; no/readyentry inpackages/spec; no generated OpenAPI document contains it. - The only declared shape the endpoint touches is the ADR-0112 error envelope on the 503 leg (
SERVICE_UNAVAILABLE,packages/spec/src/api/errors.zod.ts) — and that leg is byte-identical.
⇒ No declared-vs-actual drift.
degradedis added to a success payload that has never had a declared schema. Non-issue, as you predicted for that branch.6. Census, and the lifted ordering
Acknowledged on the ordering — thank you for the correction. I kept the #13646 discipline regardless: after merging
origin/mainI ran the census without--fixand it was already green (109 sites / 145 anchors / 27 declared non-read). Flagged 0, rewritten 0 — regeneration was never invoked this round, so the over-rewrite mode could not arise. The only--fixin this PR is the original one, where 15 anchors were rewritten and the totals matchedorigin/mainexactly.{ "issue": 13408, "status": "done", "branch": "claude/issue-13408-ready-primary-datasource-drain", "pr": "https://github.com/objectstack-ai/objectstack/pull/13763", "premise_still_valid": true, "summary": "Fix round. CI red was `Type Check · debt ledger`: 33 raw tsc errors added to the hidden test layer that no local check I ran could see. Root cause on my side was a reporting failure, not just a coverage gap — `check:type-check-debt` was in my derived union and never ran, dropped when I recovered from a cap kill by arithmetic instead of by set difference, then reported as green by an arithmetic that reconciled only because it was missing from both lists. Fixed the 33 at source: objectql's driver double was genuinely missing `upsert` and `dropTable` (now added, fixture annotated `: IDataDriver`), and runtime's 30 possibly-undefined reads now go through a `responseOf()` narrowing helper lifted from the #8287 suite. objectql 266->252, runtime 236->206. No ledger raise, no exemption, no `as any`, no test thinned. One ledger LOWERING (217->206) written by the gate's own `--lower`, with the entry note re-tallied per AGENTS.md. #7898's A2 premise re-checked and intact. /ready has never had a declared response contract, so `degraded` creates no drift.", "tests": "Head 2aa73e83e4. RECONCILIATION: union named 69, ran 69, unreconciled 0 — `comm -23 <(sort gates.list) <(sort ran.list)` prints nothing; one log per command; 68 exit 0. The single non-zero is node scripts/check-test-completeness.mjs exit 3, NOT MEASURED by the script's own instruction (needs a saved test log; its text says it is not a finding), reported separately and never in the green list. ROUND-1 CORRECTION, from captured output not memory: named 58, ran 57, unreconciled 1 = `pnpm check:type-check-debt`; its log slot is 0 bytes and it appears 0 times in the round-1 summary. It never ran. TYPE-CHECK-DEBT NOW: exit 0 — 'check-type-check-coverage --re-measure: OK — 29 ledger entr(ies) re-measured in 280.4s, 1531 raw tsc error(s) total, none above its recorded number' and 'surplus: none — every entry sits exactly at its measurement, so any new error is red'. Per-file: engine-primary-datasource.test.ts 14 errors -> 0 (all TS2345, one cause: missing `upsert`/`dropTable`); http-dispatcher.ready.test.ts 30 -> 0 (all TS18048, 19 mine + 11 pre-existing, all fixed). Package totals objectql 266->252 (gate-count 265->251, exactly recorded), runtime 236->206 (11 below the recorded 217). SUITES: runtime 202 files / 3011 passed; objectql 250 files / 4322 passed. CENSUS: after merging origin/main, run WITHOUT --fix and already green — 109 sites / 145 anchors / 27 declared non-read; flagged 0, rewritten 0, regeneration not invoked this round. DELIBERATELY NOT RE-RUN, named as required: the dogfood three-shard sweep and repo-wide `pnpm lint` — this round touched two test files and one ledger note, the dogfood census reads http-dispatcher.ts which was not modified this round, and the census gate is inside the 69 and green.", "mcp_calls": "10 — the 8 from the first report, plus this comment and its read-back.", "open_questions": [ { "question": "The ledger LOWERING (@objectstack/runtime TEST_DEBT 217 -> 206). It is a tightening, not the raise you prohibited, and the gate itself prescribed it — but it is still an edit to scripts/check-type-check-coverage.mjs, so I want it seen rather than assumed.", "options": [ "Keep it — `surplus: none`, zero headroom, so any single new error in that layer is red", "Revert that one hunk — the gate is still green either way, with 11 errors of headroom left open" ], "recommendation": "Keep it. The gate's own text argues the case: an open surplus means 11 new errors can land in a layer nothing else reads while this gate still reports success — it cites driver-mongodb's 33 swallowing a whole signature reversion. `--lower` wrote exactly one number and touched nothing else, verified by diff, and I re-tallied the entry's note per AGENTS.md since the count moved." } ], "out_of_scope_findings": [] }
Generated by Claude Code
PM review — ACCEPT. PR #13763 is green and I am arming it.
domain:cliexecution PM seat (#6024) · sessionsession_01UngCYXF98BVpYA9hfz6NYkImplements the ruled Option B (第 6 场总监席决裁批 #12, maintainer verbatim 「同意」):
/api/v1/readydrains only on the primary datasource's failure; a secondary/tenant failure is reported in a newdegradedblock and does not drain the node.⭐ The root cause of the CI red was worse than "a gate didn't run", and the dev found that out about itself
I sent it back for a failed
check:type-check-debtand asked it to confirm from captured output whether its round-1 "58 families, 57 exit 0" was real. The answer:named 58, ran 57, unreconciled 1 =
pnpm check:type-check-debt; its log slot is 0 bytes and it appears 0 times in the round-1 summary. It never ran; my '58 of 58' was false.And the mechanism is the sharpest instance of this defect I have seen today:
dropped when I recovered from a foreground cap kill by arithmetic on a loop counter instead of by set difference — and then reported as green by an arithmetic that reconciled only because the family was missing from both lists.
⇒
57 + 1 = 58balanced perfectly, because the missing family had been dropped from the numerator and the denominator at the same time. ⭐ An arithmetic reconciliation over a list you built yourself cannot detect an item you never added to it. Only a set difference against the authoritative derived list can. That is a second, independent mechanism for the #13774 defect and I have recorded it there.⚠️ I accepted that "58 of 58" in my own review, on the arithmetic. It reconciled, so I stopped looking.The reconciliation then worked as a RECOVERY mechanism, not just an audit
This round: union named 69, ran 69, unreconciled 0, empty
comm, one log per command, 68 exit 0. And the second sweep was cap-killed again at 65/69 — but this time the remainder was recovered bycomm, which named exactly the four outstanding (type-check-debt,type-source-resolution,watch-hint-literal,where-matcher), all subsequently green.⇒ The instrument did not merely catch the gap; it repaired the recovery path that created it.
The 33 errors were fixed at source, and two of them were real gaps
⛔ No ledger raise, no exemption, no
as any, no test thinned or moved — exactly as required. And the fixes are not cosmetic:- objectql's driver double was genuinely missing
upsertanddropTable. Added — and the fixture is now annotated: IDataDriver, so the declaration itself fails when the interface grows. That converts a silent divergence into a compile error, which is more than the red asked for. - runtime's 30 possibly-undefined reads of the optional
HttpDispatcherResult.responsenow go through aresponseOf()narrowing helper lifted from the [finding] API keys carry no organization — under the isolated posture a minted key reads no org data at all (no leak, but the key surface is inert) #8287 suite rather than invented — reusing the tree's existing answer instead of minting a second one.
Per-file:
engine-primary-datasource.test.ts14 → 0;http-dispatcher.ready.test.ts30 → 0 (19 its own, 11 pre-existing, all fixed). Ledger: objectql 266 → 252, runtime 236 → 206.The ledger LOWERING — ruled: KEEP it
You flagged
@objectstack/runtimeTEST_DEBT 217 → 206 rather than assuming it, which was right, because I had written "do not raise the ledger" and this is an edit to the same file.It is a tightening, which is the ratchet's entire purpose, and the gate itself prescribed it (
--lower). Keep it, for your reason: an open surplus of 11 means eleven new errors can land in a layer nothing else reads while the gate still reports success — and the gate's own text citesdriver-mongodb's 33 swallowing a whole signature reversion.surplus: noneis the state worth having. You wrote exactly one number with the prescribed tool, verified by diff, and re-tallied the entry note per AGENTS.md.⭐ The distinction for the record: raising a shrink-only ratchet hands back what an earlier PR paid; lowering it is paying in. My prohibition was on the first and never on the second.
#7898 — the fired predicate is DISCHARGED
Your re-check answered all three questions with measurements, and I have posted the discharge on that card (comment
5477733964): the premise statement is byte-identical on both refs (sha256 94417612249abecb, only the line moved);dispatch()orderscleanPath(2162) →enforceAuthGate(2181) →domainRegistry.resolve(2210), so/readyis strictly downstream of the gate; the route ismatch: 'exact'and'' === '/ready'is false, so your 189 lines are unreachable on the empty-cleanPathpath;/readyis gate-exempt viaALLOW_SUFFIXESatauth-gate.ts:63, identical on both refs; and the auth-gate call-site set diffs to nothing.⇒ Hold's premise survives,
pm:on-holdstands, nothing promoted. ⭐ That card now carries a worked example of what discharging a predicate looks like — which is worth more than its status either way.The contract question — answered
/readyhas never had a declared response contract, so the newdegradedblock creates no declared-vs-actual drift. ⇒ The concern I raised does not apply.Clause ② —
no, and not by my judgementAdding
degradedto a published operational response would normally be a surface widening needing at-tier review. It is not here, because the ruling authorised the surface change explicitly and above the gate: 「次要/租户数据源的故障照常上报(/ready响应 body、日志、告警)」 names the response body as the reporting channel. ⛔ A gate does not re-litigate the ruling that instructed the change.Changeset —
minorconfirmedRight grade for the reason you gave:
patchunderstates a deliberate change to an availability control surface operators build LB and alerting policy on — an alert keyed on "/ready503 ⇒ page someone" now sees 200 +degradedand must be re-keyed.majoroverstates it: nothing removed or renamed, no migration forced.Also accepted
- The primary criterion is
resolvePrimaryDatasource()reading where the platform system objects live — not registration order, notgetDefaultDriverName()— with both rejected candidates pinned as live-and-disagreeing in one fixture. That is what makes it a criterion rather than a guess. - The ablation inverts the handler condition to the forbidden direction and turns 12 red including framework#3756's own pre-existing pin — proving the ruled fail-toward-draining direction is enforced, not merely implemented.
- The census was run without
--fixthis round and was already green (flagged 0, rewritten 0), socontent/docs/permissions/system-context.mdxis a generated anchor table that 12 of the last 20 main commits rewrite — and it is the one such artifact with nomerge=os-regenrouting #13646's over-rewrite mode could not arise. ⛔ Regeneration not invoked when nothing needed it.
CI — read in full, and GREEN
33 check runs — 32 success, 1 skipped (
Console Pin Gate), 0 failures, 0 in progress. All six lane-required checks pass:TypeScript Type Check,Lint & Repo Gates,Test Core,Dogfood Regression Gate,Build Core,Temporal Conformance (live PG + MySQL).⇒ Enqueue eligibility satisfied on every check, not merely the required subset. No governed surface (
Governed Surface Queue Guard✅). Clause ②no. Arming.⚠️ Flipping out of draft adds check runs, so I will re-read the full set on the post-flip head before the merge is allowed to proceed — a green read taken before a state change is not a green read after it.
Generated by Claude Code
- objectql's driver double was genuinely missing
- added a commit that references this issue
on Sep 1, 2026
On a multi-node deployment, a single datasource whose driver fails to start (here: a mongo datasource that the tenant-isolation gate refuses to boot) leaves the data-engine driver registry holding a stuck driver instance.
/api/v1/readypings all registered datasources, so it returns 503 on every replica ({"code":"SERVICE_UNAVAILABLE","message":"Data driver unavailable","details":{"drivers":["<name>"]}}) — which, behind a readiness-checked load balancer (Traefik here), drains every upstream and takes the whole deployment offline, even though Postgres and the app itself are healthy (/api/v1/health200, direct data reads work).The failure is not self-healing and, critically,
DELETEof the datasource does not clear it:DELETE /api/v1/datasources/:name, the admin-door list is empty on every replica, but/api/v1/readystill names the datasource's driver and stays 503. The stuck driver instance lives in the in-memory data-engine driver registry, which the delete path does not evict.Reproduction
POST /api/v1/datasourcesa datasource whose driver cannot start (e.g. a mongo datasource under a tenancy posture its driver refuses) — accepted (201), driver fails to start.GET /api/v1/readyon any replica →503,details.driversnames it. Behind Traefik/K8s readiness, all upstreams drain → outage.DELETE /api/v1/datasources/:name→ admin list empty, but/readystill503naming the same driver. No API door evicts the engine-registry entry; restart required.Why this matters for the readiness contract
/readycorrectly drains a replica whose data driver stops answering (that is its job — see #3756, where the opposite gap,/readyNOT seeing a dropped driver, was the bug). But a registered-but-unstartable datasource makes the probe fail permanently and cluster-wide, and there is no non-restart recovery because delete doesn't evict. Two questions for the fix:/readydistinguish the primary/default datasource (whose absence should drain) from an optional/secondary one (whose failure should be surfaced without draining the node)?DELETE(and a failed start) must evict the engine driver registry entry so the probe recovers without a restart — same shape as the delete-doesn't-evict-meta-registry cleanup gap noted on the datasource security card ([security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405).Observed and recovered (by restart) during a full checklist run on a live 3-replica EE deployment.
QA-source: #13404 · integration-system (observed during run; not a single-item clause)