Repository navigation
[finding] driver-registry eviction is per-replica — the driver registry has no cluster propagation in either direction, so a deleted datasource still drains the replicas that did not serve the DELETE #13805
Description
Activity
- addedbugSomething isn't workingSomething isn't workingpriority:p1High: required for production / M2High: required for production / M2and removed
on Aug 31, 2026 分诊定级 →
domain:services· p1 · bug ·pm:queue。摘finding。锚定。 要发出集群事件的写路径在
packages/services/service-datasource/src⇒domain:services。⚠️ packages/objectql的驱动注册表是申报的跨域肢(domain:engine),由 services 席在认领评论申报文件面 —— ⛔ 但注意卡片指出的层次问题:packages/objectql依赖一条 cluster bus 本身就是一个分层问题,若最终形状要求那条边,回报而不是自行添加。定级 p1
#13578 之前:一个驱动起不来的 datasource 让
/api/v1/ready在全部副本上 503,Traefik 上游全部抽干,只有重启能清(QA 运行 #13404 把它列为 availability 提取项)。#13578 之后:服务了那次 DELETE 的那一个副本恢复,其余 N-1 个继续挂着卡死的驱动直到重启。
⇒ 从「全挂」变成「N-1 挂」。这仍然是一次生产事故,只是小了一点 —— 在三副本部署上,删掉一个坏 datasource 之后仍有 2/3 的上游是 503。与 #13578 同判 p1。
⭐ 卡片纠正了一处会误导实现的框架,这是它最重要的贡献
这不是 #13405 / #13609 的那种不对称。 那两张说的是
/api/v1/meta/datasource的元数据注册表 —— create 全集群广播、delete 不镜像,是一条本来就有广播路径的注册表上的真不对称。
驱动注册表是对称的:两半都不广播。 只给 delete 加广播,会让 delete 比 create 更懂集群 —— 那是反方向的新不对称,不是修复。⇒ ⛔ 派单令必须写死这一条。 #13578 的卡把 #13405 的框架带了过来,而它不适用;一个照着「补上 delete 的广播」去做的 dev 会造出一个新的不对称,而且看起来像修好了。
三个待决,全部在 services 席权限内 —— 除了第 2 条的分层问题
- 哪个方向。 push(datasource 写入时发
cluster作用域事件,create 与 delete 都发以保持对称)vs pull(各副本周期性或按探针与共享 datasource 记录对账)。⭐ 卡片给的判据很好:pull 不需要新传输且漏消息后能自愈;push 恢复更快。 ⇒ 在多副本 EE 部署上,「漏消息后能自愈」这条权重不低 —— 本轮已有 cli: serve's cluster-driver load registers into the CJS registry while the ESM Runtime reads the ESM one — OS_CLUSTER_DRIVER=redis silently downgrades to "not registered" (post-#10645) #13330 证明该环境里模块实例、注册表状态都可能出岔。 - 走谁的通道。
EventScopeSchema已声明cluster作用域与 at-least-once 投递,@objectstack/service-cluster有传输实现 —— 但 datasource 路径上没有任何东西接到它。⚠️ 这条如果要求objectql依赖 cluster bus,升级回报。 - 幂等。
unregisterDriver返回true/false且可安全重复 ⇒ 重复投递已经无害;create 的重放不明显安全 —— 这一条要在实现前答,⛔ 不要假设对称。
⚠️ 与 #13330 的交叉:那张卡记录了@objectstack/service-cluster的driverRegistry是模块级单例、不挂 globalThis,在 CJS/ESM 双构建下注册与查找不在同一张表上。⇒ 若本卡走 push 且经 cluster bus,它会跑在那条已知有缺陷的路径上。⛔ 派单前确认 #13330 的处置状态;两张若同时在飞,交给同一个 dev。Refs:#13578(驱逐原语与按副本修复;其 PR 声明了本卡这一剩余)· #13609 / #13405(元数据注册表的传播缺口,形状不同)· #13404(QA 来源,availability 提取项)· #13330(cluster 注册表的双实例缺陷)。
Generated by Claude Code
- 哪个方向。 push(datasource 写入时发
转入决策箱 ——
pm:queue→needs-user-decision(domain:servicesPM 席,座位贴 #6021)⛔ 本席不自裁,理由是卡自己写的:「This is design surface, not a defect fix, which is why #13578 filed it instead of building it」,并且三个待决问题里有一个是分层问题 ——
packages/objectql是否可以依赖一条集群总线。那不是一张 p1 缺陷卡能顺手带过的形状,派下去等于让 dev 替维护者选架构。⚠️ 同时记一句反向的话,免得这张卡被读成「被搁置」:p1 定级本席不下调。#13578 之后的实际状态是「服务 DELETE 的那个副本恢复,其余副本继续端着卡死的 driver 直到重启」—— 比「重启每个副本」好,离「不用重启」还差一截,而这是运行时可达的降级态。本席的建议(供裁决,⛔ 不是裁决)
先做 pull,⛔ 不做 push —— 且两半都要,不许只做 delete 一半。
- ⭐ 对称性是这张卡最硬的一条,卡自己已经论证过,采纳:driver 注册表两个方向都不广播,所以只为 delete 加广播会造出一个反方向的新不对称,那是把问题换个方向放,不是修复。⇒ ⛔ 任何只覆盖 delete 的方案直接出局。
- pull 优于 push,理由是幂等与自愈:卡已实测
unregisterDriver返回true/false、重复调用安全 ⇒ 重复投递无害;而 create 的重放不显然安全(卡自己指出的)。push 模型正好把风险压在这个不安全的一侧,pull 模型则天然只做收敛。加上 pull 不需要新传输、丢消息后自愈,分层问题(问题 2)也一并绕开 ——packages/objectql不必知道集群总线存在。 - push 唯一买到的是恢复更快。 ⇒ 这是个可量的取舍,不是信仰问题:请在裁决时说明「一个被删数据源在其余副本上继续卡死多久是可接受的」。有这句话,pull 的对账周期就是导出来的,不是拍的;没有这句话,任何一方都没法验收。
派发上的现成边界(裁完即可直接下派,⛔ 不必再走一轮分诊)
- ⛔ 不碰 [finding] QA observed /api/v1/meta/datasource serving a deleted datasource's entry cluster-wide, but MetadataManager's unregister DOES broadcast — the observed prolongation sits in an unidentified seam #13609 / [security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405 的元数据注册表。卡已经把这个混淆挡掉了:那是一个有广播路径、只缺 delete 镜像的真不对称;driver 注册表是两边都没有。
⚠️ datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578 的卡当初把那套框架搬了过来,是搬错的。 - ⛔ 不碰 datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578 已落地的驱逐原语本身。
- 定级 p1 保留,车道
domain:services保留。
Generated by Claude Code
Unlock scan — released to
pm:queue; both release double-checks passeddomain:servicesseat, sessionsession_01AUF1NoViznQK32gqpK8wS8(seat post #6021). Surfaced by patrol row H19.- Condition from the MOST RECENT conversion comment (the 2026-09-01T00:46Z ruling record, not any earlier state): "sequenced after runtime: TS-config boot registers a 'metadata' service without attachClusterPubSub — cross-node invalidation disabled; new object gives OBJECT_NOT_FOUND on non-writing replicas, never heals #13331-A lands … returns to the queue with its dispatch shape pre-written". Measured: runtime: TS-config boot registers a 'metadata' service without attachClusterPubSub — cross-node invalidation disabled; new object gives OBJECT_NOT_FOUND on non-writing replicas, never heals #13331 closed 2026-09-01T14:07:42Z via merged PR feat(metadata-protocol,service-cluster): fan runtime metadata mutations out to peer replicas #14183 — the cluster-invalidation bridge family this card is ruled to adopt is on
main. Condition met. - No merged PR on this card newer than the conversion comment — zero closing PRs reference it.
Label swap done in one write (
pm:blocked→pm:queue), read back and verified. The maintainer's ruling stands unchanged and will be quoted verbatim in the dispatch order: symmetric create+delete signals through the #13331-A bridge family, convergence by re-reading shared datasource records, no second propagation mechanism.Deliberately NOT dispatched in this batch (deferral notes, so the next dispatch inherits them):
- automation/approvals: 多副本集群下审批流每级节点(除首级)被重复创建 —— approve 后恢复读到滞后一拍的流运行态,同级要批两次(单副本零重复) #13617 is being dispatched this batch and its fix may also land in
service-cluster's bridge/registration surface — batch-independence rule says serialize when two cards can plausibly touch the same package. - PR fix(service-cluster): stop the metadata bridge reporting “bridged” over an in-process cluster bus #14228 is in the merge queue touching
metadata-cluster-bridge-plugin.tsandcontent/docs/kernel/cluster.mdx— branch after it lands, or treat those files as read-only exemplars. - The dual-build registry split triage flagged as a hazard on this path (cli: serve's cluster-driver load registers into the CJS registry while the ESM Runtime reads the ESM one — OS_CLUSTER_DRIVER=redis silently downgrades to "not registered" (post-#10645) #13330) is FIXED on
mainvia merged PR fix(types,cli): resolve host-declared packages through theimportcondition, and read the cluster registry instead of assuming it #14042 — the emit path no longer risks registering into a registry the runtime does not read. Re-verify rather than inherit at dispatch time.
Generated by Claude Code
- Condition from the MOST RECENT conversion comment (the 2026-09-01T00:46Z ruling record, not any earlier state): "sequenced after runtime: TS-config boot registers a 'metadata' service without attachClusterPubSub — cross-node invalidation disabled; new object gives OBJECT_NOT_FOUND on non-writing replicas, never heals #13331-A lands … returns to the queue with its dispatch shape pre-written". Measured: runtime: TS-config boot registers a 'metadata' service without attachClusterPubSub — cross-node invalidation disabled; new object gives OBJECT_NOT_FOUND on non-writing replicas, never heals #13331 closed 2026-09-01T14:07:42Z via merged PR feat(metadata-protocol,service-cluster): fan runtime metadata mutations out to peer replicas #14183 — the cluster-invalidation bridge family this card is ruled to adopt is on
Claim: PM loop round 1 (services seat, cap raised to 4 by the maintainer 2026-09-02)
Session:session_01AUF1NoViznQK32gqpK8wS8
Branch:claude/issue-13805-driver-registry-cluster-convergence
Worktree:objectstack-issue-13805
Domain:domain:services
File surface:packages/services/service-datasource/src/**(the write paths:datasource-admin-service.tsregister/unregister/reregister seams at ~:816/:836/:844 anddatasource-admin-plugin.ts~:405/:411 atd63c8a25) andpackages/services/service-cluster/src/**(a new adopter of the #13331-A bridge family, beside the two exemplars); ⛔packages/objectql/**is READ-ONLY — a layering change there is stop-and-report, exactly as the ruling and triage both say (stop on breach; explain in the report)
Container & model: M,mode:subagent,model: fable (claude-fable-5)— dispatch-gates atd63c8a25prints "no path-derived mandate: the surface hits none of the 3 declared glob(s)"; the tier is set by the Clause-② mechanical floor: an adopter plugin exported fromservice-cluster(the exemplars are exported) is a new exported symbol ⇒yesregardless of back-compat
Clause-②: yes
Serial constraints cleared:service-clusteris free (PR #14228 MERGED 01:11Z; no open PR touches it orservice-datasource— verified against the open-PR list). In-flight siblings: #13617 holdsplugin-approvals/service-automationand is fenced from WRITINGservice-cluster(its claim), so the earlier fold-or-serial answer flips from serial to parallel by construction; #13926 holdsservice-analytics. The dual-registry hazard triage flagged (#13330) is FIXED onmainvia merged PR #14042. Read-coupled: #13609/#13405 (metadata registry — ⛔ different sink, not touched), #13578 (eviction primitive — ⛔ not touched).
Generated by Claude Code
Claim (dev subagent, dispatched under the PM's round-1 claim above)
Session:session_01AUF1NoViznQK32gqpK8wS8(the dispatching PM session; this comment is the dev seat's own claim marker, scratchpad keyissue-13805-knife1)
Branch:claude/issue-13805-driver-registry-cluster-convergence(pushed empty as the write-route probe)
Worktree:objectstack-issue-13805, base909a4417(origin/main at claim time; the PM's readings were taken atd63c8a25and are being re-derived on this base)
File surface (unchanged from the PM claim):packages/services/service-datasource/src/**(emit seams + the convergence read) andpackages/services/service-cluster/src/**(the #13331-A bridge family adopter);packages/objectql/**READ-ONLY. Docs:content/docs/kernel/cluster.mdxonly if the landed channel needs a sentence besidemetadata.mutated.
Ruling and triage constraints acknowledged verbatim: symmetric create+delete signal through the same bridge shape, convergence by re-reading the shared datasource records, no second propagation mechanism; the/api/v1/meta/datasourcemetadata registry (#13609/#13405) and #13578's eviction primitive are not touched.Generated by Claude Code
Generated by Claude Code
- added a commit that references this issue
on Sep 2, 2026 os-dev-report
{ "issue": 13805, "status": "done", "branch": "claude/issue-13805-driver-registry-cluster-convergence", "pr": "https://github.com/objectstack-ai/objectstack/pull/14347", "premise_still_valid": true, "summary": "Implemented the ruled shape: DatasourceAdminService now publishes an address-only signal {originNode, name} on a new cluster channel datasource.mutated after createDatasource, updateDatasource and removeDatasource (symmetric; migrateCredential deliberately not), and every peer converges its live pool from its OWN read of the durable sys_metadata row through the pool seams it already owns (build / rebuild-in-place keeping the old pool on failure / evict via the #13578 door / leave a matching pool alone) — a new optional convergePool seam on DatasourceAdminServiceConfig that DatasourceAdminServicePlugin supplies, with the record each live pool was built from remembered by the pool seams so duplicates are no-ops and stray names cannot touch a code-defined pool. The attach seam attachDatasourceMutationPubSub(pubsub, nodeId) mirrors the protocol's (idempotent pair, originNode loopback, per-name ordered receipt); only IPubSub from spec/contracts crosses it, service-datasource takes no cluster dependency and objectql is handed no bus (dependency direction measured: service-cluster depends on core+spec only; neither package depends on the other). Adopter placement deviates from the PM's suggested route in one way, within the ruling: a THIRD LANE in MetadataClusterBridgePlugin (how #13331-A itself adopted the family) rather than a fourth plugin — no Runtime wiring change, no new export from service-cluster, same isInProcessClusterDriver guard. PM assumptions re-derived at 909a4417 (origin/main moved from d63c8a25 during dispatch): service-datasource/src had zero cluster occurrences apart from a mongodb URL in a test fixture; the #13330 dual-registry fix is on main; the emit seams sit where the claim said. Not touched: the /api/v1/meta/datasource metadata registry (#13609/#13405 — convergence deliberately does not write the peer's metadata registry, pinned) and #13578's primitive. Platform note, not corrected: the PR body read back with a second session footer and a --- appended by the platform after create. Label needs:contract-review hung on the PR and read back from the labels endpoint.", "tests": "Base 909a4417; final head 6be669f6 (the union below was run at 6be669f6 after the last commit). Dependency closure built first under the shared lock: os-verify-lock -c \"pnpm --filter '@objectstack/service-datasource^...' build\" → VERDICT command-exit 0 (held 243s, waited 355s); then both packages built (VERDICT command-exit 0). pnpm --filter @objectstack/service-datasource typecheck → exit 0, 0 error TS; tsc --listFiles counts datasource-cluster-convergence.test.ts once (tests are in its program). pnpm --filter @objectstack/service-cluster exec tsc --noEmit → 1 error, src/memory/memory.contract.test.ts:26 TS2322, pre-existing (git diff origin/main on that dir empty; the package's DEBT ledger entry says errors: 1, code-tier TS2322) — holds, nothing added. Tests under the lock at 6be669f6: os-verify-lock -c 'pnpm --filter @objectstack/service-cluster exec vitest run --maxWorkers=2 && pnpm --filter @objectstack/service-datasource exec vitest run --maxWorkers=2' → 'Test Files 7 passed (7) / Tests 105 passed (105)' and 'Test Files 30 passed (30) / Tests 635 passed (635)', 'VERDICT command-exit 0 · held the lock 25s · waited 486s'. New two-node rig datasource-cluster-convergence.test.ts: 21 cases (Arm B: DELETE on writer evicts on peer, create builds on peer from its own read, connectivity change rebuilds in place, active flips; Arm A controls with no attach: DELETE leaves the peer's driver, create leaves the peer pool-less; address-only payload for all three doors; duplicate delivery no-op by identity; label-only no churn; ordered create-then-update; stray name leaves a code-defined pool; peer rebuild failure keeps the old pool; unreadable row not spent as gone; publish failure never fails the write; detach; late replica converges by rehydration; loopback; malformed payload; idempotent attach; no-convergePool host; migrateCredential silent). metadata-cluster-bridge-plugin.test.ts +8 lane-3 cases (cross-process attach with the verbatim 'bridged datasource.mutated → cluster.pubsub (node=node-a)' line; memory driver skips at debug; absent/bare/no-cluster skips; throwing attach reported without taking lanes 1-2 down; shutdown detaches lane 3 even when lane 2's detach throws). ABLATION (committed tree, direction declared in the test header; no rebuild needed and none done — the test imports the package's own src, no dist on the path): perl -0pi removed the publishDatasourceMutation call from removeDatasource; mutation confirmed on disk by grep -c (call sites 2 → 1, marker ABLATION-13805 0 → 1); lock run of the convergence file → 'Tests 2 failed | 19 passed (21)', VERDICT command-exit 1, the reds being exactly 'Arm B — DELETE on the writer evicts the driver on the PEER' and 'publishes name + originNode only … symmetric'; restore via git checkout HEAD -- ABSOLUTE_PATH inside trap EXIT INT TERM, proved by git hash-object == HEAD blob (bedd9a61…) and git diff HEAD empty; git status clean afterwards. First test run (pre-fix tree 7401705d) had 1 red: my stray-signal case asserted no warn at all while the plugin's own boot-time 'datasource nav contribution skipped' warn fired in the harness — assertion narrowed to the receive path's warn in 6be669f6, not a code change. Gates from dispatch-gates.mjs --commands (42 commands / 35 families, run unlocked, exit captured before any pipe): all exit 0 except four NOT MEASURED by their own exit-3/prerequisite text — check:type-check-debt and check:dual-build-cjs-loads (need the whole workspace built; declared narrowing, CI runs them), check-test-completeness (parses a CI vitest summary), scripts/pm/check-half-states.mjs (PM patrol, network, timed out). engine-double-contract: the new file's findOne/update/delete doubles call the metadata-core asserts; ledger learned them via its own --write (+3 rows, 0 lost), bare re-run green. check:nul-bytes OK (7825 files). ESLint narrowed: pnpm exec eslint --no-inline-config --format json over the 6 changed TS files → 6 files, 0 errors, 0 warnings; population evidence: the config's only global ignores are NEVER_LINTED (node_modules/dist/build/.next/.turbo); invariance: eslint.config.mjs states it never enables type-aware linting for any file, so the diff moves no untouched file's verdict. Full pnpm lint not run (repo-level, CI-owned).", "mcp_calls": "2 — both search_issues for the finding dedup (control query hit #13805 → search live this session; docs-gap query returned no existing card). Claim comment, issue read (page payload + REST comments), PR create, label add, read-backs, finding creation and this report all went through git / zero-quota payload / REST (session gate open: repo-scoped GET 200, POST 201).", "open_questions": [], "out_of_scope_findings": [ "filed as #14339: docs — content/docs/kernel/cluster.mdx describes the cluster-invalidation family as metadata.changed only; the metadata.mutated (#13331) and datasource.mutated (#13805) lanes are undocumented, and the bridge plugin's docblock points at §5 where the fan-out lives in §6 (finding label, unassigned; one docs-only change for the whole family)" ] }Generated by Claude Code
Generated by Claude Code
Collected — PR #14347 under in-seat Clause-② review; carriers paired
domain:servicesseat, sessionsession_01AUF1NoViznQK32gqpK8wS8.Carrier fix:
needs:contract-reviewwas on PR #14347 only; hung on this card in one write, read back present;check-clause2-carriers --pair 14347reads both carriers as agreeing. Clause-② stands as declared:yes— two new exported symbols from@objectstack/service-datasource(DATASOURCE_MUTATION_CLUSTER_CHANNEL,ClusterDatasourceMutationPayload, re-exported fromindex.ts), two public methods and one optional config key on already-exported types.Verified against the tree, not the report (merge-base
909a4417→ head6be669f6): 8 files —service-datasource(admin service, admin plugin,index.ts, a new 21-case two-node convergence test),service-cluster(metadata-cluster-bridge-plugin.ts+ its test: a third lane, not a new plugin —service-clusterexports nothing new), one changeset (minorfor both packages, per the mechanical floor), andscripts/engine-double-contract.pinned.json(+3 rows learned via the gate's own--writefor the new test'sfindOne/update/deletedoubles).packages/objectql/**untouched (read-only fence held).service-datasource/package.jsongains noservice-clusterdependency — the layering line triage drew holds; onlyIPubSubfrom spec contracts crosses the seam.content/docs/releases/untouched.Ruling conformance at first read: symmetric emit at create/update/delete (
migrateCredentialdeliberately silent); receipt converges by re-reading the sharedsys_metadatarow, not the receiving replica's metadata registry; no poll loop or second channel. The maintainer's ruled direction is implemented as stated; the adopter placement (a lane in the existing bridge plugin, as #13331-A itself did) is within the ruling and adds no export toservice-cluster.Pending before any landing action: the isolated contract-review verdict at the contract-review tier (adversarial second opinion, transcript-verified) and every check green on head
6be669f6(CI's Lint & Repo Gates is the authoritative lint reading — the dev ran ESLint narrowed to the six changed files with population evidence, not the repo-wide scan). PR stays draft until both. Follow-up filed by the dev: #14339 (docs,cluster.mdxnames onlymetadata.changed).
Generated by Claude Code
In-seat Clause-② contract review — isolated second opinion, adopted verbatim (PM seat, session_01AUF1NoViznQK32gqpK8wS8)
Provenance: the reviewer was a context-isolated subagent fed only this card (body and every comment), the rulings quoted on it, and PR #14347 itself at head
6be669f6— not the dispatch order and not the seat's own reading. Transcript tier verification: 47 harness-stampedmodelfields, 47 of 47 readclaude-fable-5-1(CONTRACT_REVIEW_TIERisclaude-fable-5; the served-5-1tier is treated as at-tier per finding #14303). Verdict PASS ⇒ adopted verbatim below, unedited. In this same strokeneeds:contract-reviewis stripped from both carriers (this card and PR #14347) per the maintainer's 2026-08-31 ruling that a PASS is cleared and landed by the dispatching seat. Landing follows once every check on6be669f6is green: ready → auto-merge into the merge queue. The reviewer's four non-blocking follow-ups are recorded here for triage; none is required for merge.--- verdict, verbatim ---
Clause-② second-opinion verdict — PR #14347 (head
6be669f6, base909a4417, closes #13805)Read: issue #13805 body + all 7 comments (ruling 2026-09-01T00:46Z os-sam; triage 2026-08-31T14:37Z os-warren), PR body, full diff per file (8 files, +1312/−4), exemplars
authz-cluster-bridge-plugin.ts/metadata-cluster-bridge-plugin.tsat the merge-base, the protocol'sattachMetadataMutationPubSub/publishMetadataMutation(packages/metadata-protocol/src/protocol.ts:5031-5105),DatasourceConnectionServiceconnect/disconnect/reconnect/recordState/disconnectAll,datasourceConnectivityChanged,IPubSub/PublishOptions.partitionKey, the #13398 transcript on PR #13592,.changeset/config.json, the shipped secret binder.1. Ruling conformance
- Emit symmetric — yes.
publishDatasourceMutation(name)is the last statement of all three record doors indatasource-admin-service.ts:createDatasource(afterputDatasourceRecord+tryRegisterPool),updateDatasource(unconditional, after persist + the datasource update never rebuilds its driver —already-registeredshort-circuits the reconfigure path, so the OLD pool stays live and the admin UI reports success #13804 pool decision),removeDatasource(afterdeleteDatasourceRecord+tryRemoveSecret+tryUnregisterPool). One channeldatasource.mutated, one payload{ originNode, name }. Not a delete-only broadcast. migrateCredentialdoes not publish. Justified as pool-neutral on every replica, but it does change connectivity-bearing fields of the shared row (config.<key>removed,external.credentialsRefadded — both compared bydatasourceConnectivityChanged). Peers'livePoolRecordstherefore go stale until the next signal of any kind, which then rebuilds the peer's pool — a one-time churn on e.g. a label-only edit after a migration. Not a ruling breach; see §6(c).- Receipt converges from the SHARED durable record — yes.
convergePool→loadDatasourceRow(engineOf(), name)→engine.findOne('sys_metadata', { type:'datasource', name, state:'active' })→applyConversionsToStoredItem. It never callsmetadataOf().get(...)(the per-replica registry). The payload carries no verb and no content; test "publishes name + originNode only" pinsObject.keys(payload).sort() === ['name','originNode']for all three doors, so nothing is replayed. - No second mechanism — confirmed.
grepof the src diff: nosetInterval/setTimeout/poll; onesubscribecall; onepublishcall site.convergeChainsis a per-name promise chain for receipt ordering, not a scheduler. - No new dependency edge — confirmed. No
package.jsonin the diff.datasource-admin-service.tsadds onlyimport type { IPubSub } from '@objectstack/spec/contracts'(spec is an existing dep); the plugin adds an intra-package import; the bridge imports nothing new and reaches the seam asunknownvia duck-typing.packages/objectql/**untouched;service-clusterdeps staycore+spec.
2. Family shape (lane 3 vs lane 2, site by site)
Site Lane 2 (exemplar) Lane 3 (PR) getServicethrowsdebug debug seam absent debug debug isInProcessClusterDriverdebug + return, before .calldebug + return, before .call— same position (lookup → seam → in-process → attach)attach ok info bridged metadata.mutated → cluster.pubsub (node=…)info bridged datasource.mutated → cluster.pubsub (node=…)(pinned verbatim)attach throws error mutation-lane attach failederror datasource-lane attach failedshutdown own try/catch, error mutation-lane detach error, field clearedown try/catch, error datasource-lane detach error,detachDatasource = undefined, sequenced after lane 2's block (test pins a throwing lane-2 detach does not strand it; second shutdown detaches once)Service-side seam is a structural copy of
attachMetadataMutationPubSub(protocol.ts:5031):(pubsub, nodeId)pair-idempotency returning the disposer, detach-before-reattach,originNodeloopback, malformed payload dropped, fire-and-forget apply withwarn,partitionKey: datasource:<name>mirroring${type}:${name}. Differences: uses the injectedLogger(info?/debug?/warn) instead ofconsole.*; adds onedebugfor a host withoutconvergePool.Loggerinlogger.tsis unchanged anderroris never invoked inservice-datasource; the onlyerrorcalls are the bridge's kernel-logger sites at exactly the positions the exemplar already uses. Ruling (c) satisfied — no published sink shape gains or raises a level.3. Convergence safety
Property Test ( datasource-cluster-convergence.test.ts)Code path Verdict Duplicate delivery no-op "a duplicate delivery is a no-op" — same instance, created1,evicted0matching branch → registerPool(row)→connect→attemptConnectgetDriverByNamehit →already-registered, no buildHolds for the driver registry. Caveat: connect()then callsrecordState(connection-service.ts:475-477) and overwrites the peer's ownconnectedverdict withalready-registered;disconnectAll()(:863-866) filtersstatus==='connected', so that pool is skipped at graceful teardown and the admin list on that replica showsalready-registered. Pre-existing on the serving replica (label-only edit →tryRegisterPool), now triggered on every peer by every duplicate/label-only signal. Not pinned by the test. Non-blocking.Stray name touches nothing "a code-defined pool survives a stray name" — code driver same instance, not closed, evicted0,created0, noconvergingwarnrowundefined,liveundefined → returnHolds. livePoolRecordsis populated only by this plugin'sregisterPool/reregisterPool; callers are the service's doors,rehydratePools(filteredorigin==='runtime' && active),convergePool. Caveat: it records names attempted, not built — a runtime record colliding with an engine-registered code driver that has nodatasourcemetadata item enters it viaalready-registered; a later delete signal woulddisconnect(name)that code driver. The serving replica's ownremoveDatasource→tryUnregisterPool→disconnectdoes the same with no gate at all, so mirrored, not new.Rebuild failure keeps old pool "a rebuild that fails on the peer keeps the peer's OLD pool" — identity, connected, not closed, defschemaMode:'external'restoredreregisterPool(live,row)→reconnect→restoreOldPool()withprevious = liveHolds. Note livePoolRecords.set(next)runs beforereconnect, so after a failed rebuild it names a config the pool does not serve — identical to the serving replica's record store after #13804's failed rebuild; only cost is one extra correct rebuild if the row later reverts.Unreadable row ≠ gone "NOT spent as gone" — identity, not closed, evicted0, warnconverging 'analytics' after a peer write failedloadDatasourceRowthrows on JSON parse → chain.catch→ warnHolds for parse failure; a findOnerejection takes the same throw path structurally (not separately pinned). Gap:loadDatasourceRowreturnsundefinedwhen!engine?.findOne(data service unresolvable throughsafeGetService) →!row→ evictslive. "Engine absent" is conflated with "row absent"; the invention rule the PR cites for parse failures is not applied here. Reachable only ifgetService('data')throws afterkernel:ready(shutdown ordering while the bus still delivers). Untested. Non-blocking; one-line hardening: throw or return without touching the pool when the engine is absent.Publish failure never fails the write "a publish failure never fails the write" — summary returned, pool connected, warn void publish(...).catch(warn)Holds. Nit: a synchronous throw from publishwould escape (no try around the call) — the protocol exemplar has the identical shape, so family-consistent.Create chased by update "two signals in quick succession converge in order" — final /tmp/new.db,drivers.size===1enqueueConvergenceper-name chainTerminal state pinned. The "every pool opened is live or closed" claim is asserted as expect(stale).toBeGreaterThanOrEqual(0)— vacuous. Bus order is sequential (create awaited before update); whether convergence #1 is still in flight when #2 lands is timing-dependent, so the chain is not deterministically exercised. Weak pin, not a defect.Peer evicting a pool it did not build: only via the
livePoolRecords"attempted" case above (mirrors the serving replica). Evicting on a transient read error: only via the engine-absent path above.4. Semver / Clause-②
From my own
git diff -U0 | grep export:export const DATASOURCE_MUTATION_CLUSTER_CHANNEL,export interface ClusterDatasourceMutationPayload, plus bothindex.tsre-exports. Public surface additions:DatasourceAdminService.attachDatasourceMutationPubSub(),.detachDatasourceMutationPubSub(),DatasourceAdminServiceConfig.convergePool?. Published payload keys:originNode,name; channel literal'datasource.mutated'. Private additions:clusterPubSub,clusterNodeId,clusterUnsubscribe,convergeChains,enqueueConvergence,publishDatasourceMutation; pluginlivePoolRecords,convergePool, module-privateloadDatasourceRow; bridgedetachDatasource,attachDatasourceLane. PR body's list matches.@objectstack/service-datasource: minor— required by the floor, declared. Correct.@objectstack/service-cluster: minor— no new export, so the floor does not require minor; it is over-declared relative to the floor but defensible for newfeatbehaviour in an exported plugin. It does not matter: both packages are in the singlefixedgroup in.changeset/config.json, so the group bumps by the highest declared level (minorfromservice-datasource) either way.- No
content/docs/releases/edit.changeset-no-majorunaffected.
5. Tests
- Real boots:
makeReplicaconstructs andinit()s +start()s the realDatasourceAdminServicePlugin, which builds the realDatasourceAdminServiceand realDatasourceConnectionService;serviceis captured fromregisterService('datasource-admin', …), which the plugin feedsthis.service(the instance, not a facade — so the bridge's duck-type on the production object is valid). Fakes sit at the correct boundaries: data engine (routing throughmetadata-coredouble asserts; +3 pinned rows inengine-double-contract.pinned.json), driver factory, metadata service, bus. The seam under test is not mocked. - Arm A controls:
attach:false→ nosubscribe,published.length===0; the peer keeps its boot-rehydrated driver (unclosed) after the writer's DELETE and stays pool-less after create. Both replicas sharestore.rowsbut hold separatedriversmaps, so Arm A proves that store sharing alone does not move the peer's registry — Arm B's convergence is the bridge's doing. - Ablation consistency: removing the publish from
removeDatasourcereddens exactly (i) "Arm B — DELETE on the writer evicts the driver on the PEER" and (ii) "publishes name + originNode only … symmetric" (published.length2 ≠ 3). I checked every other case: none depends on the delete publish (after detachexpects 0,migrateCredentialexpects 0, the active flip uses update). The declared direction matches the structure. - Bridge tests mock the seam — appropriate for the bridge, whose contract is the duck-typed attach; the verbatim info line is pinned.
- Weak pins: ordering (vacuous
stale>=0); duplicate no-op does not assert the connection-state ledger; engine-absent path untested.
6. Boundary flags
- (a) Peer metadata registry not converged — load-bearing. Declared out of scope per ruling/triage ([finding] QA observed /api/v1/meta/datasource serving a deleted datasource's entry cluster-wide, but MetadataManager's unregister DOES broadcast — the observed prolongation sits in an unidentified seam #13609/[security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405) and pinned as a non-write. Consequence the PR body does not state: on the host-config boot,
getDatasourceRecordreadsmetadataOf().get('datasource', name)(per-replica). A datasource created at runtime through replica A now has a live pool on B and C (new, from this PR) but no metadata record there until restart or [finding] QA observed /api/v1/meta/datasource serving a deleted datasource's entry cluster-wide, but MetadataManager's unregister DOES broadcast — the observed prolongation sits in an unidentified seam #13609. A DELETE routed to B fails onremoveDatasource's first line (not found) — no delete, no signal, no eviction anywhere. So "a deleted datasource stops draining every replica" holds when the DELETE reaches a replica whose metadata registry has the record (the creator, or any replica after a restart). Pre-PR, B could not serve the DELETE either but had no pool to be stuck on; post-PR, B has the pool, can be drained by it, and cannot serve the DELETE that clears it. Compliant with the ruling; must be recorded on [finding] QA observed /api/v1/meta/datasource serving a deleted datasource's entry cluster-wide, but MetadataManager's unregister DOES broadcast — the observed prolongation sits in an unidentified seam #13609 as the dependency that completes the ruled outcome. - (b) Credential rewrap-in-place is invisible to convergence. On a host whose
SecretBinder.bindreturns the same ref (the case the baseupdateDatasourcecomment names), a rotation changes nothing in the shared row, so peers keep the old credential until a connectivity-bearing edit or restart. The shippedcreateDatasourceSecretBindermints a newhandle.idperbind(new ref →externaldiffers → peer rebuilds), so the shipped host converges. A documented limit of the address-only shape, not a defect. - (c)
migrateCredentialnon-publish → stalelivePoolRecordson peers → one-time rebuild on the next signal of any kind (§1). - (d) At-most-once delivery → loss bound is the next boot; declared in the changeset and pinned by the "late replica" case.
- (e)
content/docs/kernel/cluster.mdxuntouched ([finding] docs: content/docs/kernel/cluster.mdx describes the cluster-invalidation family asmetadata.changedonly — themetadata.mutated(#13331) anddatasource.mutated(#13805) lanes are undocumented #14339 filed, unassigned); the new docblock still points at §5.
7. VERDICT: PASS
Strongest reason: the diff implements exactly the ruled shape and nothing beside it — one symmetric address-only signal at all three record-write doors, received into a convergence that re-reads the shared
sys_metadatarow through the plugin's own pool seams, guarded so only names this plugin registered can be evicted, with no second mechanism and no new dependency edge; the two-real-plugin-over-one-store rig with Arm A controls and a structurally consistent ablation pins the ruled outcome rather than the harness.Non-blocking follow-ups the author should take (none required for merge):
loadDatasourceRow: treat an absent engine as "read did not complete" (throw / no-op), not as "row gone".convergePoolmatching branch: skipregisterPoolwhenthis.connection?.getState(name)?.status === 'connected'so a duplicate/label-only signal does not rewrite the peer's verdict toalready-registered(which excludes the pool fromdisconnectAllat teardown); keep the call when the state is failed so the retry semantics survive.- Replace the vacuous
expect(stale).toBeGreaterThanOrEqual(0)with an assertion that every non-live driver the factory produced is closed. - Record §6(a) on [finding] QA observed /api/v1/meta/datasource serving a deleted datasource's entry cluster-wide, but MetadataManager's unregister DOES broadcast — the observed prolongation sits in an unidentified seam #13609 so the "DELETE must reach the creating replica" limit is tracked as the remaining half.
REVIEWER_MODEL_SELF_REPORT: Claude Fable 5.1 (claude-fable-5-1)
--- end verdict ---
Generated by Claude Code
- Emit symmetric — yes.
Erratum to the verbatim verdict above (read-back after posting): the platform sanitizer removed two angle-bracket tokens from the reviewer's text. Restored here in escaped spelling so they survive; nothing else in the verdict differs from the reviewer's output.
- §1, second bullet:
config.should readconfig.<key>(the credential key removed from the shared row bymigrateCredential). - §2, paragraph after the table:
partitionKey: datasource:should readpartitionKey: datasource:<name>.
Generated by Claude Code
- §1, second bullet:
- added a commit that references this issue
on Sep 2, 2026
Found while implementing #13578. That card fixed the eviction itself; this is the propagation half it explicitly declared out of scope rather than improvising.
The state today, measured
The ObjectQL driver registry has no cluster propagation in either direction:
packages/services/service-datasource/srcpublishes one.DatasourceAdminPlugin.rehydratePools->registerPoolper record).So after #13578,
DELETE /api/v1/datasources/:namerecovers/api/v1/readyon the replica that served it, and the other replicas keep the stuck driver until they restart. Better than "restart every replica", still not "no restart".Why this is NOT the #13405 / #13609 asymmetry
Worth stating, because the #13578 card carried that framing across and it does not transfer. #13405 (and #13609) describe the
/api/v1/meta/datasourcemetadata registry, where create broadcasts cluster-wide and delete does not mirror it — a genuine asymmetry on a registry that HAS a broadcast path.The driver registry is symmetric: neither half broadcasts. Adding a broadcast for delete alone would make delete more cluster-aware than create, which is a new asymmetry in the opposite direction rather than a repair.
What a fix needs to decide
This is design surface, not a defect fix, which is why #13578 filed it instead of building it:
cluster-scoped event on datasource writes, both create and delete, so the two stay symmetric) or a pull model (each replica reconciles its registry against the shared datasource records periodically, or on a probe). The pull model needs no new transport and is self-healing after a missed message; the push model recovers faster.EventScopeSchemaalready declaresclusterscope withat-least-oncedelivery, and@objectstack/service-clusterimplements transports — but nothing on the datasource path is wired to it, andpackages/objectqltaking a dependency on a cluster bus is a layering question in its own right.unregisterDriverreturnstrue/falseand is safe to repeat, so a duplicate delivery is already harmless; a create replay is not obviously so.Refs
Generated by Claude Code