Repository navigation
[finding] QA observed /api/v1/meta/datasource serving a deleted datasource's entry cluster-wide, but MetadataManager's unregister DOES broadcast — the observed prolongation sits in an unidentified seam #13609
Description
Activity
- addedbugSomething isn't workingSomething isn't workingpriority:p2Medium: important, M3Medium: important, M3and removed
on Aug 31, 2026 分诊定级 →
domain:engine· p2 · bug ·pm:queue。摘finding。锚定。 四个候选接缝里三个落
domain:engine(MetadataManager 的unregister/notifyWatchers、list-cache TTL、以及 #13578 那类 registry 驱逐);restoreRuntimeDatasources在packages/runtime(domain:cli)是唯一的例外。主车道按多数与根因指向定domain:engine,并与它的姊妹卡 #13578(已pm:dispatched,domain:engine,p1)一致。⭐ 入队而不是留在 finding 箱 —— 因为交付物是一次已具名的测量,不是一个待猜的方向
本卡是「观察 vs 源码互相矛盾」型,这类卡默认该留箱等测量。但本卡不同:它把四个候选接缝逐一点名了,每一个都可判定:
- 部署形态里 pubsub 根本没挂(广播发了,没人听);
restoreRuntimeDatasources在驱逐后又把条目播回去;- list-cache TTL 供着陈旧条目 —— 已关的 集群对端的元数据写入不失效本节点的
listCache/ registry —— 收到广播的节点最长 30s 继续服务旧定义 #5109 精确测过这个形状(集群 peer 上可达 ~30s),⚠️ 但本次观察到的时长读起来比一个 TTL 长; - 与 datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578 同类,只是隔一个 registry —— 若 DELETE 漏掉一个 registry,它可能漏掉不止一个。
⇒ 这是一份可执行的排除清单,派给一个能起多节点部署的 dev 就能收敛。⛔ 不需要先裁决。
⭐ 与 #13578 同批做,理由是第 4 条候选
#13578(data-engine 驱动 registry 在 DELETE 时不驱逐,
/ready继续点它的名)已经在飞。本卡的第 4 条候选说的正是「同一个 DELETE 路径可能漏掉多个 registry」。⇒ 交给同一个 dev 最省:它已经在读那条生命周期路径,多验一个 registry 的成本远小于第二个人从头建多节点环境。派单时请点名这个配对。定级 p2 —— 曝露那一半已经关了,剩下的是生命周期
原始 QA 归因是 P0/security(凭据曝露被延长)。但读侧曝露已由 PR #13604 关闭(嵌套位置在每一扇读门上都已脱敏)。所以本卡剩下的是:哪个接缝在供一个已被删除的条目、供多久 —— 正确性与生命周期问题,不再是曝露问题。⇒ p2,不是 p1。
⚠️ 但不要因此把它读软:一个 DELETE 之后仍被三个副本供着的条目,本身就是错误答案,与它是不是凭据无关。⛔ 两条纪律,卡片立得很好,原样保留
- 观察与反证都不丢弃。 QA 当时的读法是「create 广播,delete 不镜像」,源码证明这在 MetadataManager 层不成立(
unregister同样经CLUSTER_CHANNEL扇出,收端跑invalidateForForeignWrite)。⇒ 朴素机制被证伪,但被观察到的延长是真的。⛔ 不许用「源码看起来是对的」结案 —— 那是拿源码读数去否定现场读数。 - 复现细节按 skills: five principle gaps the #9296 wave exposed in checklist-test / checklist-author / dogfood-verification #9387 披露裁定扣留,本卡与任何提取卡上都不publish。
Refs:#13405(母卡 + 带反证的 dev 报告)· #13578(在飞姊妹卡)· #5109(已关,listCache TTL 先例)· #13404(QA 来源)。
Generated by Claude Code
zhuangjianguo commented
on Aug 31, 2026 CollaboratorMore actionsCLAIMED + dispatch order — #13609
- Session:
session_01F3jdziLbAPGeceVNmSox5L· Branch:claude/issue-13609-meta-datasource-stale-entry· Worktree:../objectstack-13609-meta-datasource, offorigin/main - Labels:
pm:queue→pm:dispatched - Clause ②: does NOT fire on the measurement.
⚠️ If your reading turns into a fix that changes what a read door serves, re-declare from what you actually built. ⛔ Upward only.
⚠️ On triage's pairing request — honoured in substance, not in formTriage asked for this to go to #13578's dev, since it is already reading that lifecycle path ("多验一个 registry 的成本远小于第二个人从头建多节点环境"). That seat has since finished — it delivered PR #13829 and closed out. Rather than resume a seat carrying ~356k tokens of a different card, I am handing you its findings directly, below, so you inherit the context without re-deriving it. ⛔ Do not re-measure what §A already establishes — cite it and move on.
§A — Inherited from #13578 / PR #13829, measured, ⛔ do not re-derive
The sibling card is done and its measurements bear directly on your candidate 4:
- The DRIVER registry (data-engine) had no eviction door at all —
this.drivershad exactly one.setsite and zero.deletesites repo-wide. PR Give the driver registry an eviction door, so a deleted datasource stops draining /ready #13829 addsunregisterDriverand wires it into datasource DELETE, kernel teardown, engine teardown and failed-start rollback. - ⭐ The DRIVER registry has NO cluster broadcast in EITHER direction — no datasource create or delete emits a cluster event; each replica populates its own registry at boot from the shared records via
rehydratePools. So its per-replica behaviour is symmetric, not an asymmetry. - ⭐⭐ Give the driver registry an eviction door, so a deleted datasource stops draining /ready #13829 explicitly distinguishes itself from your card: the create-broadcasts/delete-doesn't asymmetry [security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405 records is on the
/api/v1/meta/datasourceMETADATA registry — "a different registry with a different propagation story." That metadata registry is yours.
⇒ A three-way tension you must resolve, and it is the heart of this card:
- QA observed: the deleted entry kept being served on all three replicas.
- [security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405's reading: "create broadcasts, delete does not mirror."
- Source counter-evidence:
MetadataManager.unregisterdoes fan out onCLUSTER_CHANNELvianotifyWatchers, and the receiving peer runsinvalidateForForeignWrite— the same fan-outregisteruses.
All three cannot be simultaneously true as stated. Finding which one is wrong, or which unstated condition reconciles them, is the deliverable.
Zone 1 — binding. ⛔ NOT re-adjudicable.
1.1 — ⛔⛔ NEITHER the observation NOR the counter-evidence may be discarded. Triage states it as a discipline and I am restating it because it is the way this card fails:
⛔ 不许用「源码看起来是对的」结案 —— 那是拿源码读数去否定现场读数。
The naive mechanism is falsified. The observed prolongation was real. ⛔ "The source looks correct, therefore QA was mistaken" is a forbidden conclusion. If you cannot find the seam, say you cannot find it — that is a legitimate outcome; explaining the observation away is not.
1.2 — This is a MEASUREMENT card, not a fix card yet. Triage: "这是一份可执行的排除清单" — an executable elimination list. Your deliverable is which seam, and for how long. ⛔ If a fix becomes obvious and small you may propose it, but the measurement is what is owed.
1.3 — ⛔ Reproduction details are WITHHELD under the #9387 disclosure ruling. Do not publish them on this card or anywhere else. Work from the source seams.
1.4 — The read-side exposure is ALREADY CLOSED by PR #13604 (nested positions redacted on every read door). ⛔ Do not re-open or re-litigate that half.
⚠️ But do not therefore read this card as soft: an entry still served by three replicas after a DELETE is a wrong answer regardless of whether it carries a credential.
Zone 2 — the four candidate seams. Each is decidable; ⭐ report a verdict on EVERY one.
Triage named these and they are your elimination list. ⭐ Three of the four are statically determinable from source — do those first, and do not let the fourth's difficulty block them.
A2.1 — pubsub not attached in the deployment shape QA used (broadcast emitted, nobody listening). ⭐ Partly static: is pubsub wired by default, and what happens to
notifyWatcherswhen it is not? A no-op broadcast would reconcile all three observations at once — this is my prime suspect and the cheapest to check.A2.2 —
restoreRuntimeDatasourcesre-seeding the entry after eviction. ⭐ Statically readable.⚠️ It lives inpackages/runtime(domain:cliterritory) — you may READ it; if the fix lands there, that is a routing question, so report rather than edit.A2.3 — list-cache TTL serving the stale entry. ⭐ Statically readable, with a precedent: closed #5109 measured exactly this shape for cluster peers at up to ~30s.
⚠️ The card's own caveat is load-bearing: the QA prolongation reads LONGER than a TTL. So if you land on TTL, you must explain the duration mismatch or the answer is incomplete.A2.4 — the same class as #13578, one registry over. ⇒ "若 DELETE 漏掉一个 registry,它可能漏掉不止一个." ⭐ Statically readable, and §A above already gives you half of it: the driver registry's answer is known. Ask the same question of the metadata registry — does DELETE reach its eviction door at all, and does that door broadcast?
Zone 3 — advisory
⚠️ A live multi-node deployment is probably not available to you. Triage assumed a dev who can raise one. If you cannot, ⛔ do not fake it and do not treat single-node behaviour as evidence about cluster behaviour — narrow the list statically, then declare the live-measurement gap as NOT MEASURED, naming which candidate it leaves open. A well-narrowed list plus an honest gap is a good outcome here.- ⭐ A zero needs a firing positive control, as always — if you find "delete does broadcast", prove your probe would have seen it not broadcasting.
- ⛔ Never edit
content/docs/releases/. ⛔ Nevergit stash. Worktree-first.
STOP conditions
- You are about to conclude the observation was mistaken on source-reading alone ⇒ that is 1.1. Report the contradiction unresolved instead.
- The seam is in
packages/runtime'srestoreRuntimeDatasources(A2.2) ⇒ different lane; report, do not edit. - Answering requires a live multi-node deployment you cannot raise ⇒ narrow statically, declare the gap, stop. ⛔ Do not simulate one and present it as the measurement.
- Any other open PR claims your paths.
⚠️ PR Give the driver registry an eviction door, so a deleted datasource stops draining /ready #13829 is OPEN and unmerged, touchingpackages/objectql/src/engine.ts(driver-registry region ~5215/~7850) andpackages/services/service-datasource/src/datasource-connection-service.ts. PR Refuse an undeclared field a before-hook writes — the post-hook half of the declared-field door, one envelope on every driver #13864 is OPEN, touchingpackages/objectql/src/engine.tsat ~9490/~10470,packages/runtime,packages/lint. Two more seats are live: [finding]If-Match: ""silently DISABLES optimistic concurrency — a quoted-empty entity-tag is read as "no token" and the guarded write proceeds unguarded #13576 (packages/metadata-protocol/src/protocol.ts) and driver-memory:bulkUpdateandbulkDeleteare stillPromise.all(map(...)), so a refused row leaves the earlier rows of the batch applied — the third and fourth batch doors of the driver #13340 did not reach #13435 (packages/drivers/driver-memory). ⇒ Your likely surface (MetadataManager / meta read doors / list cache) should be disjoint from all four — if it is not, STOP and report.
Verification expected
A verdict on each of the four seams — eliminated, confirmed, or NOT MEASURED with the reason; the three-way tension in §A resolved or explicitly left open; a firing positive control on any zero; and, if you identify the seam, how long it serves the stale entry. ⛔ Report anything unmeasured as NOT MEASURED, never as a pass. ⛔ No test skipped, disabled, quarantined, allow-listed, deleted or re-baselined. ⛔ Do not publish withheld reproduction detail.
Generated by Claude Code
- Session:
- added a commit that references this issue
on Aug 31, 2026 zhuangjianguo commented
on Aug 31, 2026 CollaboratorMore actionsos-dev-report
{ "issue": 13609, "status": "done", "branch": "claude/issue-13609-meta-datasource-stale-entry", "pr": "https://github.com/objectstack-ai/objectstack/pull/13883", "premise_still_valid": true, "summary": "Resolved the three-way tension statically, with an executable positive+negative control (no live multi-node deployment was available or attempted). unregister() DOES fan out via CLUSTER_CHANNEL/notifyWatchers exactly as the source counter-evidence claims -- confirmed by a positive-control test where a shared transport correctly evicts a peer immediately. The seam is one layer down: Runtime's shipped default (cluster option omitted) resolves to the in-process 'memory' cluster driver, whose own doc-comment states it has no cross-process delivery; the split-brain guard only fires when an operator explicitly declares multi-node via OS_EXPECT_MULTI_NODE/OS_CLUSTER_REPLICAS, so a default-configured multi-replica deployment silently gets N independent, non-communicating pubsub instances. A peer that never receives the metadata.changed event keeps the deleted row in its in-memory registry (no TTL there -- only listCache has one), and readListUncached() never re-checks a registry hit against the loader, so the stale entry survives indefinitely, not for ~30s -- reproduced with a second test showing it served past 10 list-cache TTL windows. Verdicts: A2.1 CONFIRMED as the seam (refined: pubsub is attached and does publish, but the shipped default transport does not cross OS processes and nothing catches an undeclared multi-node deployment on it); A2.2 ELIMINATED (restoreRuntimeDatasources runs once at boot only, reads the already-corrected DB, cannot explain steady-state cross-replica staleness without a restart -- also: it lives in packages/services/service-datasource/src/datasource-admin-plugin.ts, not packages/runtime/domain:cli as the dispatch order stated, a minor correction); A2.3 ELIMINATED as the direct cause but identified as the reason the TTL precedent looked plausible -- the stale answer comes from the untimed registry map, not the 30s listCache, which is exactly why the observed duration reads longer than any TTL; A2.4 CONFIRMED as the same symptom class as #13578 but a different mechanism -- unlike the driver registry (zero .delete() sites), this metadata registry's eviction door exists and correctly broadcasts; what fails is the cluster transport, not the wiring. Landed two vitest tests pinning both directions in packages/metadata/src/metadata-manager-cluster.test.ts, opened as a measurement-only draft PR (no changeset; skip-changeset label applied via the additive REST endpoint, 200, confirmed by the label read-back in the same response).", "tests": "pnpm --filter @objectstack/metadata exec vitest run src/metadata-manager-cluster.test.ts (via scripts/pm/os-verify-lock.sh) -> 'Test Files 1 passed (1)', 'Tests 15 passed (15)' (13 pre-existing + 2 new: the shared-transport positive control, and the separate-transport reproduction that advances vi fake timers by LIST_CACHE_TTL_MS * 10 and still finds the deleted row via both get() and list()). Ran all 24 of the 26 dispatch-gates-derived local-gate commands that could run without a full 78-package workspace build (12 pnpm check:* plus convention-triggered kind gates, and 10 direct node scripts -- check:engine-double-contract, check:where-matcher, check:test-source-alias, check:cross-package-test-inputs, check:type-check-coverage, check:objectql-double-limit, etc.) under one scripts/pm/os-verify-lock.sh slot: all EXIT=0, captured via 'cmd >> log 2>&1; ec=$?' (never a bare piped $?). The remaining 2 (pnpm check:dual-build-cjs-loads, node scripts/check-test-completeness.mjs) plus a manual probe of check:type-check-debt's --re-measure all self-report PREREQUISITE NOT MET at exit 3 (their own documented signal, distinct from a finding's exit 1) because they read built dist/*.d.ts or a saved CI turbo-test log that a local worktree does not produce without a full-farm pnpm build across 44-78 packages -- recorded NOT MEASURED, not pass or fail, matching each gate's own self-test output verbatim. Also ran pnpm exec eslint on just the touched file (clean, zero output) and a standalone tsc --noEmit -p packages/metadata sanity pass: only pre-existing TEST_DEBT-ledgered errors at line 242 and earlier (all pre-existing content, unrelated to this diff); zero new errors on the added lines (402 through 513).", "mcp_calls": "4 (issue_read get, issue_read get_comments, create_pull_request, add_issue_comment for this report)", "open_questions": [ { "question": "The confirmed seam (A2.1: shipped default in-process cluster driver + a split-brain guard that stays silent unless an operator explicitly declares multi-node) has no landed fix in this PR, only a pinning test. How should the finding be closed out?", "options": [ "A: File a targeted follow-up fix card once a remedy is chosen -- e.g. sharpen the split-brain guard to WARN (not throw) when a metadata service and a cluster service are both wired but no multi-node declaration is present, since that combination is itself informative even without a declared replica count", "B: Documentation/observability-only fix -- louder wording in the cluster option's docblock and content/docs/kernel/cluster.mdx, plus an explicit boot-time INFO log line stating the resolved driver and whether cross-process delivery actually works, so a misconfigured deployment is visible from logs alone", "C: A live multi-node verification pass with a real distributed driver, to close the one residual NOT MEASURED gap this report carries: whether the QA deployment that originally reported #13405/#13609 was in fact running the default memory driver, versus hitting a genuinely different defect in a real distributed driver" ], "recommendation": "B first -- cheap, safe, and closes the 'fails silently' gap the guard's own source comment already names. A only if the maintainer wants an enforced guard rather than a louder log. C only if B/A land and a still-open production report remains unexplained, since this measurement already accounts for every observed symptom (3 replicas, no error, duration past any TTL) without needing a live deployment." } ], "out_of_scope_findings": [] }
Generated by Claude Code
1 remaining item
zhuangjianguo commented
on Aug 31, 2026 CollaboratorMore actions#13883 landed —
0c143ececbonmain, verified by content (the file's test count moved 13 → 15; 112 insertions, one file). This card stays open and staysneeds-user-decision;closed_by_pull_requestsis empty, so nothing here was closed by a body keyword.What is now delivered
The measurement this card asked for is an executable, CI-checked artifact rather than a comment. Two tests in
packages/metadata/src/metadata-manager-cluster.test.ts:- Control —
control: WITH a working cross-node transport, unregister() on A evicts B immediately. This is the firing positive control the finding needed: it proves the probe below would have seen a working broadcast had the transport had one. - Reproduction —
WITHOUT it (the shipped default across real replicas): B keeps serving the deleted row past 10 list-cache TTL windows.
Which candidate seam it settled
Of the four seams this card named, the measurement discriminates cleanly:
Candidate seam Verdict pubsub not attached in the QA deployment shape This is it, one layer down. Runtime's shipped default (clusteromitted) resolves todefineCluster({})→driver: 'memory', andMemoryPubSub's own doc-comment says "No cross-process delivery". Every replica process builds its own instance; nothing wires them together.restoreRuntimeDatasourcesre-seedingNot the sink. (Also note: that function lives in packages/services/service-datasource, notpackages/runtime— a routing caveat attached earlier to this card was wrong and is withdrawn.)list-cache TTL (#5109 precedent) Ruled out, and this is the sharp part. The stale row is served from the in-memory registry, which carries no TTL at all — only listCachedoes — andreadListUncached()never re-checks a registry hit against the loader. That is why the QA prolongation read longer than a TTL: it is not TTL-bound at all, so it does not expire.same class as #13578 Same symptom class, opposite mechanism. #13578's driver registry had no eviction door at all. Here the door exists and does broadcast — unregister()fans out onCLUSTER_CHANNELexactly as #13405's counter-evidence says. The transport is what fails.So the observation and the counter-evidence are both vindicated: nothing in this card's contradiction needs discarding.
What still needs you
The routing decision is unchanged and still yours — whether to sharpen the split-brain guard, document the default more loudly, or leave it. One note that may reduce what you have to decide at once:
⭐ Option B (a boot-time INFO log naming the in-process default + louder docs) is separable and dispatchable without a ruling. ADR-0010 path A governs the guard — when
assertClusterDriverSafeForTopologymay fire — not whether the runtime may say out loud which cluster driver it resolved. If you want that half moving while the guard question waits, say so and I will dispatch it.The remaining open sub-question, unmeasured either way: whether the guard should fire on the shipped default when the operator has not declared multi-node. That is a deliberate ADR-0010 design choice, not a defect, which is why it is not being fixed unilaterally.
Generated by Claude Code
- Control —
⭐ The unidentified seam this card names has been identified — by #13331's reading
This card's open question is "
MetadataManager's unregister DOES broadcast — the observed prolongation sits in an unidentified seam." A measurement taken for #13331 (apriority:p0on the create side of the same mechanism) locates it. Recording here so the decision is taken once, for both.The seam: the runtime authoring path never reaches
MetadataManagerat all.packages/metadata-protocol/src/protocol.tsowns that whole path —saveMetaItem→applyRegistryWriteThrough→applyObjectRegistryMutation→engine.registry.registerObject— and contains zero references to cluster, pubsub, orgetService('metadata'). Measured with two positive controls on the same instrument (applyRegistryWriteThroughin the same file = 16 hits; the same pattern againstmetadata-manager.ts= 43), so the zero is a reading rather than a mis-spelled query.metadata.changed's only production publisher isMetadataManager.notifyWatchers(packages/metadata/src/metadata-manager.ts:2853) — and noPUT /api/v1/meta/*reaches it on any boot shape.⇒ That is why this card's observation is self-consistent and still puzzling:
unregisterreally does broadcast, because that call arrives through the manager. The runtime authoring path does not, so its mutations mutate only the localengine.registrywhilesys_metadata(shared DB) updates fleet-wide. Hence the asymmetry both cards see — metadata reads are correct everywhere, the data plane is correct only on the replica that served the write.Why the two cards should be decided together
#13331's dev escalated rather than implementing, because the repair mints new public cross-node contract (a channel, a payload, an attach entry point) spanning
domain:engine+domain:services. That contract is exactly what this card needs too. ⇒ Designing it from the create side and again from the delete side would design the same channel twice, and the two designs would have to be reconciled afterwards — the expensive order.Their recommended shape (mirroring the in-tree
AuthzClusterBridgePluginprecedent — the state owner exposes attach-pubsub,service-clusterlate-binds atkernel:ready, and peers converge from their ownsys_metadataread rather than trusting a payload) covers registration and deregistration symmetrically, because both are mutations through the sameonMetadataMutationchoke point.⚠️ Stated as a cross-link, not a re-grading of this card and not a ruling — #13331 isneeds-user-decisionwith the full reading and a four-lens analysis in its PM review. Whoever answers that decision should pick this card up in the same breath.
Generated by Claude Code
- added a commit that references this issue
on Sep 1, 2026 os-support-ai commented
on Sep 1, 2026 CollaboratorMore actions🔓 Blocker landed — #13331's fix is on
origin/main.pm:blocked→pm:queue.The maintainer's ruling (director batch A, 2026-09-01, 「同意。」) made this card the re-verification carrier: "when #13331's fix lands, re-measure the deleted-datasource prolongation on a multi-node deployment … Close only on that measurement."
PR #14183 merged as
1403d943and was verified by content on the tree, not by the API'smergedfield — landing record. So the condition on the ruling is met and this card is now dispatchable.
⭐ Before unblocking I re-checked this card's premise against the LANDED code — because a sibling finding could have voided it
Contract review of #14183 returned a finding that would, if it applied here, have made this re-verification pointless:
deleteMetaItem's legacy raw-engine branch performs a real delete and never emits — no local listener event, and therefore nometadata.mutatedpublish, since the cluster half ridesemitMetadataMutation. A deleted entry down that branch would keep prolonging on peers exactly as this card describes, after the fix.⛔ I did not transmit that as a caveat, because it is decidable and I had not decided it. Measured on
origin/main@1403d943:deleteMetaItem(packages/metadata-protocol/src/protocol.ts:19788) forks onconst useRepoPath = overlayAllowedForRepoDel || runtimeCreateAllowedForRepoDel;
and the
datasourceentry inDEFAULT_METADATA_TYPE_REGISTRY(packages/spec/src/kernel/metadata-plugin.zod.ts:930) reads, verbatim:type: 'datasource', supportsOverlay: false, allowOrgOverride: false, allowRuntimeCreate: true,
⇒
useRepoPath = false || true = **true**⇒ a datasource DELETE takes the repository exit at:20050, and that exit emits at:20043(state: 'deleted', scoped toorganizationId).⭐ Positive control on the flag reading: across the registry
allowRuntimeCreateistrue30× andfalse9×. The field is not uniformly true, sodatasource: trueis a reading rather than an artifact of a probe that would have said "true" whatever it hit.⇒ This card's premise survives the landing. The delete direction does now fan out, and the re-verification is a genuine open question rather than a formality. (The non-emitting branch is real; it lands on types with both flags false —
api,field— and has been appended to #14179, where it belongs, not here.)
What is now owed on this card
The four-seam checklist from triage still governs, but three of the four are already settled by PR #13883's measurement (
packages/metadata/src/metadata-manager-cluster.test.ts). What #13331's landing changes is the first one:seam status going into re-verification pubsub not attached / no cross-process delivery ⭐ the one to re-measure. #13883 pinned that the shipped default resolves to the in-process memorydriver with no cross-process delivery. #14183 adds the publisher and theservice-clusterbridge — so the question is now "with a real distributed driver and the new bridge attached, does the deleted entry still prolong?"restoreRuntimeDatasourcesre-seedingELIMINATED (#13883) — boot-only, reads the corrected DB list-cache TTL (#5109) ELIMINATED as the cause, but ⚠️ the ruling flags it as the one bounded residue that may legitimately remain. A ≲30s window after a DELETE is not this card's defect; an unbounded one is.same class as #13578 ELIMINATED — same symptom, opposite mechanism (that registry had no eviction door; this one's door exists and broadcasts) ⭐ The duration is the discriminator, and it is what makes this measurable at all. #13883 established that the stale row is served from the in-memory registry, which carries no TTL — only
listCachedoes, andreadListUncached()never re-checks a registry hit against the loader. That is precisely why the original QA prolongation read longer than any TTL. So the re-verification has a sharp pass/fail: bounded (≲ onelistCacheTTL) = fixed; unbounded = not fixed. ⛔ "It cleared eventually" is not a reading unless the bound is stated.⚠️ The standing constraint on this card has NOT changedA live multi-node deployment is still what the ruling asks for, and a dev seat may not be able to raise one. The dispatch rules from the previous round still hold and I will restate them when this goes out:
- ⛔ Do not treat single-node behaviour as evidence about cluster behaviour, and do not simulate a cluster and present it as the measurement.
- If the live deployment cannot be raised: narrow statically against the landed publisher/bridge, then declare the gap as NOT MEASURED, naming what it leaves open. A well-narrowed list plus an honest gap is a good outcome here — it was last time.
- ⭐ The residual unmeasured sub-question from test(metadata): pin default
memorycluster driver's cross-process isolation (#13609 measurement) #13883 is still open and is now more answerable, not less: was the QA deployment that reported [security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405/[finding] QA observed /api/v1/meta/datasource serving a deleted datasource's entry cluster-wide, but MetadataManager's unregister DOES broadcast — the observed prolongation sits in an unidentified seam #13609 actually on the default memory driver, or did it hit a second defect in a real distributed driver? With the bridge landed, a real-driver run distinguishes these for the first time. - ⛔ Reproduction detail stays withheld under the skills: five principle gaps the #9296 wave exposed in checklist-test / checklist-author / dogfood-verification #9387 disclosure ruling.
- ⛔ Neither the observation nor the counter-evidence may be discarded (Zone 1.1, still binding). "The fix looks correct, therefore it is fixed" is the same forbidden move as "the source looks correct, therefore QA was mistaken."
Close only on that measurement, per the ruling — not on this landing.
Generated by Claude Code
os-project-manager commented
on Sep 3, 2026 CollaboratorMore actionsMaintainer ruling recorded — A: the
metadata.mutatedreceipt path bumps the local write epoch (epoch.bump('remote'), the same call the authz bridge makes), so a peer's overlay cache cannot re-hydrate a deleted datasource into the registry the bridge just healedDirector seat (objectstack #12708), summon #10, session
session_01ShyhexkB2d1AeRZ85tgAAe, 2026-09-03.Provenance (who / verbatim / where): maintainer, live PM chat with the director seat, replying to decision batch #15 in which this card was item 2 with the recommendation A (fallback B: targeted invalidation; C: hydrate on miss only; D: document only), the same recommendation the
domain:engineseat attached in its four-facet block (5505583013). Verbatim reply: 「同意」 — A is adopted as recommended.Why this card is ruled a second time. The maintainer's earlier ruling (director batch A, 2026-09-01, 「同意。」, 5486839204) folded this card into #13331 on the premise that #13331's publisher-plus-bridge fix would close the observation, and closed it only on a measurement. PR #14431 (merged
a98b61b3e) took that measurement and it reads NOT FIXED: the bridge heals the peer's registry, but one read of the peer's own/api/v1/meta/datasourcedoor inside the 30 s overlay-cache window writes the deleted row straight back, because the receipt path never moves the write epoch while the siblingauthz-invalidation-bridge.ts:71does. The earlier ruling's exit condition holds; this ruling chooses the repair.Ruled: A. On receiving a remote metadata mutation the engine bumps its local write epoch, exactly as the authz bridge does on the same substrate — one existing spelling, no new invalidation door, no change to the overlay cache's design claim that hits and misses are indistinguishable downstream. Cost, accepted: every remote metadata change makes this engine re-read its full cached row set once, the price the authz bridge already pays.
Not taken: B (a targeted-invalidation door in
metadata-protocol— a second invalidation idiom and a cross-module coupling), C (hydrate only on miss — overturns the cache's core design claim), D.Execution:
domain:enginelane,Fixes #13609. One bump on the receipt path (inapplyRemoteMetadataMutationor the cluster bridge plugin, the authz bridge as the template — the dev measures which). The two ⛔ UNBOUNDED pins inprotocol.datasource-delete-prolongation.test.tsare inverted to bounded (read-in-window ⇒ bounded within one hop), never deleted; the other arms stay as they are.patchchangesets for the packages touched. Clause-② no.packages/metadata-protocol/src/protocol.tsis a hot file: serial behind #14078's implementation. Live multi-node deployment stays NOT MEASURED, declared as before; the QA cluster-driver question from PR #13883 stays named.State transition, same stroke:
needs-user-decision→pm:queue;bug,priority:p2,domain:engineretained. Ledger: objectstack director seat post #12708, summon #10.
Generated by Claude Code
os-project-manager commented
on Sep 3, 2026 CollaboratorMore actionsMaintainer ruling recorded — A′: ruling A stands, amended on the site, the spelling, and a third pin (2026-09-03)
Provenance. Ruled by the maintainer (pm@objectstack.ai) in the director seat's live session (
session_01ShyhexkB2d1AeRZ85tgAAe), decision batch #22, 2026-09-03 after 15:35Z. Verbatim reply: 「同意」. The batch was presented with this card's recommendation as A′ (B, the literal ruling, cannot land because the pin asserts the number; C, loosening the arm to<=, is the phantom-check shape), so 「同意」 adopts A′. Ledger: objectstack#12708, batch #22 ledger comment.What stands. The 02:25Z ruling A (comment 5519392217): the
metadata.mutatedreceipt path bumps the local write epoch. Confirmed by measurement (comment 5525867467): the bump at the receipt choke point turns both ⛔ UNBOUNDED arms into bounded-at-0 ms.What is amended.
- Site.
packages/metadata-protocol/src/protocol.ts, insideapplyRemoteMetadataMutation, after the registry-convergence branch and beforenotifyMutationListenersLocal(the 集群对端的元数据写入不失效本节点的listCache/ registry —— 收到广播的节点最长 30s 继续服务旧定义 #5109 invalidate-before-notify rule). The "or the cluster bridge plugin" alternative is withdrawn: measured to be pure boot-time wiring with no receipt path and no route to the epoch, and the authz template keeps its bump in the state owner too. - Spelling.
metadata-protocolmust not import@objectstack/objectql, so the bump goes through a structural helper besidereadWriteEpochinmeta-overlay-cache.ts(bumpWriteEpoch(engine, 'remote')or the seat's equivalent), never a direct import of the epoch type. - Pins — three changes, authorized. The two ⛔ arms invert (four assertion-level
UNBOUNDEDreadings, two arms), and the SCOPED-kernel arm's bound moves fromTTL_MS(30 000) to0with its rationale rewritten to say why (on a scoped kernel the overlay cache is the only local source, and the bump retires it at convergence instead of letting it lapse). ⛔ It is not loosened to<=; the pin keeps asserting the number.
Fence.
protocol.tsstays fenced behind PR #14908 as the engine seat's policy. The measured non-overlap (#14908's onlyprotocol.tshunk at:7348; the receipt path at:5062–5199) permits the seat to branch now and land after #14908 merges — no race, no lifting of the fence from this seat.Execution.
domain:enginelane. Clause-②: no (per the original ruling).Fixes #13609on the PR that lands it.State.
needs-user-decision→pm:queuein the same stroke;bug,priority:p2,domain:enginekept; unassigned.
Generated by Claude Code
- Site.
- added a commit that references this issue
on Sep 9, 2026 - added a commit that references this issue
on Oct 7, 2026
Filed by the domain:spec seat at acceptance of PR #13604 (#13405), so the contested observation survives that card's closure. Observation and counter-evidence disagree; neither is discarded here.
The observation (QA, held evidence)
#13405's second-order note, from the QA run #13404 (reproduction withheld per the #9387 disclosure ruling): after datasource
DELETE, the admin/detail-door registry cleared but/api/v1/meta/datasourcekept serving the deleted entry on all three replicas, prolonging the credential exposure until the entry was scrubbed. The QA reading at the time: "create broadcasts, delete does not mirror."The counter-evidence (source, PR #13604's dev report on #13405)
At the MetadataManager layer the asymmetry does NOT exist:
unregisterfans out onCLUSTER_CHANNELvianotifyWatchers, and the receiving peer runsinvalidateForForeignWrite— the same fan-outregisteruses. So the naive mechanism the QA note implies is contradicted; the observed prolongation, which was real, must sit in a different seam.Candidate seams (named in the dev report, none measured)
restoreRuntimeDatasourcesre-seeding the entry after eviction;listCache/ registry —— 收到广播的节点最长 30s 继续服务旧定义 #5109 measured exactly this shape for cluster peers (up to ~30s), though the QA prolongation reads longer than a TTL;/readykeeps naming it) — if DELETE misses one registry it may miss more than one.What this card is
An observation-vs-source contradiction needing a measurement in a live multi-node deployment, not a fix card yet. The read-side exposure half is already closed by PR #13604 (nested positions now redacted on every read door), so what remains is the lifecycle question: which seam serves a deleted entry, for how long, and is #13578's fix the same fix.
Refs: #13405 (parent finding + dev report with the counter-evidence) · #13578 (open sibling, driver registry) · #5109 (closed, listCache TTL precedent) · #13404 (QA source).