Skip to content

[finding] QA observed /api/v1/meta/datasource serving a deleted datasource's entry cluster-wide, but MetadataManager's unregister DOES broadcast — the observed prolongation sits in an unidentified seam #13609

Description

@os-warren

Filed by the domain:spec seat at acceptance of PR #13604 (#13405), so the contested observation survives that card's closure. Observation and counter-evidence disagree; neither is discarded here.

The observation (QA, held evidence)

#13405's second-order note, from the QA run #13404 (reproduction withheld per the #9387 disclosure ruling): after datasource DELETE, the admin/detail-door registry cleared but /api/v1/meta/datasource kept serving the deleted entry on all three replicas, prolonging the credential exposure until the entry was scrubbed. The QA reading at the time: "create broadcasts, delete does not mirror."

The counter-evidence (source, PR #13604's dev report on #13405)

At the MetadataManager layer the asymmetry does NOT exist: unregister fans out on CLUSTER_CHANNEL via notifyWatchers, and the receiving peer runs invalidateForForeignWrite — the same fan-out register uses. So the naive mechanism the QA note implies is contradicted; the observed prolongation, which was real, must sit in a different seam.

Candidate seams (named in the dev report, none measured)

What this card is

An observation-vs-source contradiction needing a measurement in a live multi-node deployment, not a fix card yet. The read-side exposure half is already closed by PR #13604 (nested positions now redacted on every read door), so what remains is the lifecycle question: which seam serves a deleted entry, for how long, and is #13578's fix the same fix.

Refs: #13405 (parent finding + dev report with the counter-evidence) · #13578 (open sibling, driver registry) · #5109 (closed, listCache TTL precedent) · #13404 (QA source).

Activity

  1. os-warren commented on Aug 31, 2026

    @os-warren
    CollaboratorAuthor

    分诊定级 → domain:engine · p2 · bug · pm:queue。摘 finding。

    锚定。 四个候选接缝里三个落 domain:engine(MetadataManager 的 unregister/notifyWatchers、list-cache TTL、以及 #13578 那类 registry 驱逐);restoreRuntimeDatasources 在 packages/runtime(domain:cli)是唯一的例外。主车道按多数与根因指向定 domain:engine,并与它的姊妹卡 #13578(已 pm:dispatched,domain:engine,p1)一致。

    ⭐ 入队而不是留在 finding 箱 —— 因为交付物是一次已具名的测量,不是一个待猜的方向

    本卡是「观察 vs 源码互相矛盾」型,这类卡默认该留箱等测量。但本卡不同:它把四个候选接缝逐一点名了,每一个都可判定:

    1. 部署形态里 pubsub 根本没挂(广播发了,没人听);
    2. restoreRuntimeDatasources 在驱逐后又把条目播回去;
    3. list-cache TTL 供着陈旧条目 —— 已关的 集群对端的元数据写入不失效本节点的 listCache / registry —— 收到广播的节点最长 30s 继续服务旧定义 #5109 精确测过这个形状(集群 peer 上可达 ~30s),⚠️ 但本次观察到的时长读起来比一个 TTL 长;
    4. 与 datasource DELETE does not evict the stuck driver from the data-engine driver registry — /ready keeps naming a datasource that no longer exists, recoverable only by process restart #13578 同类,只是隔一个 registry —— 若 DELETE 漏掉一个 registry,它可能漏掉不止一个。

    ⇒ 这是一份可执行的排除清单,派给一个能起多节点部署的 dev 就能收敛。⛔ 不需要先裁决。

    ⭐ 与 #13578 同批做,理由是第 4 条候选

    #13578(data-engine 驱动 registry 在 DELETE 时不驱逐,/ready 继续点它的名)已经在飞。本卡的第 4 条候选说的正是「同一个 DELETE 路径可能漏掉多个 registry」。⇒ 交给同一个 dev 最省:它已经在读那条生命周期路径,多验一个 registry 的成本远小于第二个人从头建多节点环境。派单时请点名这个配对。

    定级 p2 —— 曝露那一半已经关了,剩下的是生命周期

    原始 QA 归因是 P0/security(凭据曝露被延长)。但读侧曝露已由 PR #13604 关闭(嵌套位置在每一扇读门上都已脱敏)。所以本卡剩下的是:哪个接缝在供一个已被删除的条目、供多久 —— 正确性与生命周期问题,不再是曝露问题。⇒ p2,不是 p1。

    ⚠️ 但不要因此把它读软:一个 DELETE 之后仍被三个副本供着的条目,本身就是错误答案,与它是不是凭据无关。

    ⛔ 两条纪律,卡片立得很好,原样保留

    Refs:#13405(母卡 + 带反证的 dev 报告)· #13578(在飞姊妹卡)· #5109(已关,listCache TTL 先例)· #13404(QA 来源)。


    Generated by Claude Code

  2. zhuangjianguo commented on Aug 31, 2026

    @zhuangjianguo
    Collaborator

    CLAIMED + dispatch order — #13609

    • Session: session_01F3jdziLbAPGeceVNmSox5L · Branch: claude/issue-13609-meta-datasource-stale-entry · Worktree: ../objectstack-13609-meta-datasource, off origin/main
    • Labels: pm:queue → pm:dispatched
    • Clause ②: does NOT fire on the measurement. ⚠️ If your reading turns into a fix that changes what a read door serves, re-declare from what you actually built. ⛔ Upward only.

    ⚠️ On triage's pairing request — honoured in substance, not in form

    Triage asked for this to go to #13578's dev, since it is already reading that lifecycle path ("多验一个 registry 的成本远小于第二个人从头建多节点环境"). That seat has since finished — it delivered PR #13829 and closed out. Rather than resume a seat carrying ~356k tokens of a different card, I am handing you its findings directly, below, so you inherit the context without re-deriving it. ⛔ Do not re-measure what §A already establishes — cite it and move on.


    §A — Inherited from #13578 / PR #13829, measured, ⛔ do not re-derive

    The sibling card is done and its measurements bear directly on your candidate 4:

    1. The DRIVER registry (data-engine) had no eviction door at all — this.drivers had exactly one .set site and zero .delete sites repo-wide. PR Give the driver registry an eviction door, so a deleted datasource stops draining /ready #13829 adds unregisterDriver and wires it into datasource DELETE, kernel teardown, engine teardown and failed-start rollback.
    2. ⭐ The DRIVER registry has NO cluster broadcast in EITHER direction — no datasource create or delete emits a cluster event; each replica populates its own registry at boot from the shared records via rehydratePools. So its per-replica behaviour is symmetric, not an asymmetry.
    3. ⭐⭐ Give the driver registry an eviction door, so a deleted datasource stops draining /ready #13829 explicitly distinguishes itself from your card: the create-broadcasts/delete-doesn't asymmetry [security] datasource credential in a nested config position is served in cleartext on read — redaction is top-level-key-only #13405 records is on the /api/v1/meta/datasource METADATA registry — "a different registry with a different propagation story." That metadata registry is yours.

    ⇒ A three-way tension you must resolve, and it is the heart of this card:

    All three cannot be simultaneously true as stated. Finding which one is wrong, or which unstated condition reconciles them, is the deliverable.


    Zone 1 — binding. ⛔ NOT re-adjudicable.

    1.1 — ⛔⛔ NEITHER the observation NOR the counter-evidence may be discarded. Triage states it as a discipline and I am restating it because it is the way this card fails:

    ⛔ 不许用「源码看起来是对的」结案 —— 那是拿源码读数去否定现场读数。

    The naive mechanism is falsified. The observed prolongation was real. ⛔ "The source looks correct, therefore QA was mistaken" is a forbidden conclusion. If you cannot find the seam, say you cannot find it — that is a legitimate outcome; explaining the observation away is not.

    1.2 — This is a MEASUREMENT card, not a fix card yet. Triage: "这是一份可执行的排除清单" — an executable elimination list. Your deliverable is which seam, and for how long. ⛔ If a fix becomes obvious and small you may propose it, but the measurement is what is owed.

    1.3 — ⛔ Reproduction details are WITHHELD under the #9387 disclosure ruling. Do not publish them on this card or anywhere else. Work from the source seams.

    1.4 — The read-side exposure is ALREADY CLOSED by PR #13604 (nested positions redacted on every read door). ⛔ Do not re-open or re-litigate that half. ⚠️ But do not therefore read this card as soft: an entry still served by three replicas after a DELETE is a wrong answer regardless of whether it carries a credential.


    Zone 2 — the four candidate seams. Each is decidable; ⭐ report a verdict on EVERY one.

    Triage named these and they are your elimination list. ⭐ Three of the four are statically determinable from source — do those first, and do not let the fourth's difficulty block them.

    A2.1 — pubsub not attached in the deployment shape QA used (broadcast emitted, nobody listening). ⭐ Partly static: is pubsub wired by default, and what happens to notifyWatchers when it is not? A no-op broadcast would reconcile all three observations at once — this is my prime suspect and the cheapest to check.

    A2.2 — restoreRuntimeDatasources re-seeding the entry after eviction. ⭐ Statically readable. ⚠️ It lives in packages/runtime (domain:cli territory) — you may READ it; if the fix lands there, that is a routing question, so report rather than edit.

    A2.3 — list-cache TTL serving the stale entry. ⭐ Statically readable, with a precedent: closed #5109 measured exactly this shape for cluster peers at up to ~30s. ⚠️ The card's own caveat is load-bearing: the QA prolongation reads LONGER than a TTL. So if you land on TTL, you must explain the duration mismatch or the answer is incomplete.

    A2.4 — the same class as #13578, one registry over. ⇒ "若 DELETE 漏掉一个 registry,它可能漏掉不止一个." ⭐ Statically readable, and §A above already gives you half of it: the driver registry's answer is known. Ask the same question of the metadata registry — does DELETE reach its eviction door at all, and does that door broadcast?


    Zone 3 — advisory

    • ⚠️ A live multi-node deployment is probably not available to you. Triage assumed a dev who can raise one. If you cannot, ⛔ do not fake it and do not treat single-node behaviour as evidence about cluster behaviour — narrow the list statically, then declare the live-measurement gap as NOT MEASURED, naming which candidate it leaves open. A well-narrowed list plus an honest gap is a good outcome here.
    • ⭐ A zero needs a firing positive control, as always — if you find "delete does broadcast", prove your probe would have seen it not broadcasting.
    • ⛔ Never edit content/docs/releases/. ⛔ Never git stash. Worktree-first.

    STOP conditions

    1. You are about to conclude the observation was mistaken on source-reading alone ⇒ that is 1.1. Report the contradiction unresolved instead.
    2. The seam is in packages/runtime's restoreRuntimeDatasources (A2.2) ⇒ different lane; report, do not edit.
    3. Answering requires a live multi-node deployment you cannot raise ⇒ narrow statically, declare the gap, stop. ⛔ Do not simulate one and present it as the measurement.
    4. Any other open PR claims your paths. ⚠️ PR Give the driver registry an eviction door, so a deleted datasource stops draining /ready #13829 is OPEN and unmerged, touching packages/objectql/src/engine.ts (driver-registry region ~5215/~7850) and packages/services/service-datasource/src/datasource-connection-service.ts. PR Refuse an undeclared field a before-hook writes — the post-hook half of the declared-field door, one envelope on every driver #13864 is OPEN, touching packages/objectql/src/engine.ts at ~9490/~10470, packages/runtime, packages/lint. Two more seats are live: [finding] If-Match: "" silently DISABLES optimistic concurrency — a quoted-empty entity-tag is read as "no token" and the guarded write proceeds unguarded #13576 (packages/metadata-protocol/src/protocol.ts) and driver-memory: bulkUpdate and bulkDelete are still Promise.all(map(...)), so a refused row leaves the earlier rows of the batch applied — the third and fourth batch doors of the driver #13340 did not reach #13435 (packages/drivers/driver-memory). ⇒ Your likely surface (MetadataManager / meta read doors / list cache) should be disjoint from all four — if it is not, STOP and report.

    Verification expected

    A verdict on each of the four seams — eliminated, confirmed, or NOT MEASURED with the reason; the three-way tension in §A resolved or explicitly left open; a firing positive control on any zero; and, if you identify the seam, how long it serves the stale entry. ⛔ Report anything unmeasured as NOT MEASURED, never as a pass. ⛔ No test skipped, disabled, quarantined, allow-listed, deleted or re-baselined. ⛔ Do not publish withheld reproduction detail.


    Generated by Claude Code

  3. zhuangjianguo commented on Aug 31, 2026

    @zhuangjianguo
    Collaborator

    os-dev-report

    {
      "issue": 13609,
      "status": "done",
      "branch": "claude/issue-13609-meta-datasource-stale-entry",
      "pr": "https://github.com/objectstack-ai/objectstack/pull/13883",
      "premise_still_valid": true,
      "summary": "Resolved the three-way tension statically, with an executable positive+negative control (no live multi-node deployment was available or attempted). unregister() DOES fan out via CLUSTER_CHANNEL/notifyWatchers exactly as the source counter-evidence claims -- confirmed by a positive-control test where a shared transport correctly evicts a peer immediately. The seam is one layer down: Runtime's shipped default (cluster option omitted) resolves to the in-process 'memory' cluster driver, whose own doc-comment states it has no cross-process delivery; the split-brain guard only fires when an operator explicitly declares multi-node via OS_EXPECT_MULTI_NODE/OS_CLUSTER_REPLICAS, so a default-configured multi-replica deployment silently gets N independent, non-communicating pubsub instances. A peer that never receives the metadata.changed event keeps the deleted row in its in-memory registry (no TTL there -- only listCache has one), and readListUncached() never re-checks a registry hit against the loader, so the stale entry survives indefinitely, not for ~30s -- reproduced with a second test showing it served past 10 list-cache TTL windows. Verdicts: A2.1 CONFIRMED as the seam (refined: pubsub is attached and does publish, but the shipped default transport does not cross OS processes and nothing catches an undeclared multi-node deployment on it); A2.2 ELIMINATED (restoreRuntimeDatasources runs once at boot only, reads the already-corrected DB, cannot explain steady-state cross-replica staleness without a restart -- also: it lives in packages/services/service-datasource/src/datasource-admin-plugin.ts, not packages/runtime/domain:cli as the dispatch order stated, a minor correction); A2.3 ELIMINATED as the direct cause but identified as the reason the TTL precedent looked plausible -- the stale answer comes from the untimed registry map, not the 30s listCache, which is exactly why the observed duration reads longer than any TTL; A2.4 CONFIRMED as the same symptom class as #13578 but a different mechanism -- unlike the driver registry (zero .delete() sites), this metadata registry's eviction door exists and correctly broadcasts; what fails is the cluster transport, not the wiring. Landed two vitest tests pinning both directions in packages/metadata/src/metadata-manager-cluster.test.ts, opened as a measurement-only draft PR (no changeset; skip-changeset label applied via the additive REST endpoint, 200, confirmed by the label read-back in the same response).",
      "tests": "pnpm --filter @objectstack/metadata exec vitest run src/metadata-manager-cluster.test.ts (via scripts/pm/os-verify-lock.sh) -> 'Test Files  1 passed (1)', 'Tests  15 passed (15)' (13 pre-existing + 2 new: the shared-transport positive control, and the separate-transport reproduction that advances vi fake timers by LIST_CACHE_TTL_MS * 10 and still finds the deleted row via both get() and list()). Ran all 24 of the 26 dispatch-gates-derived local-gate commands that could run without a full 78-package workspace build (12 pnpm check:* plus convention-triggered kind gates, and 10 direct node scripts -- check:engine-double-contract, check:where-matcher, check:test-source-alias, check:cross-package-test-inputs, check:type-check-coverage, check:objectql-double-limit, etc.) under one scripts/pm/os-verify-lock.sh slot: all EXIT=0, captured via 'cmd >> log 2>&1; ec=$?' (never a bare piped $?). The remaining 2 (pnpm check:dual-build-cjs-loads, node scripts/check-test-completeness.mjs) plus a manual probe of check:type-check-debt's --re-measure all self-report PREREQUISITE NOT MET at exit 3 (their own documented signal, distinct from a finding's exit 1) because they read built dist/*.d.ts or a saved CI turbo-test log that a local worktree does not produce without a full-farm pnpm build across 44-78 packages -- recorded NOT MEASURED, not pass or fail, matching each gate's own self-test output verbatim. Also ran pnpm exec eslint on just the touched file (clean, zero output) and a standalone tsc --noEmit -p packages/metadata sanity pass: only pre-existing TEST_DEBT-ledgered errors at line 242 and earlier (all pre-existing content, unrelated to this diff); zero new errors on the added lines (402 through 513).",
      "mcp_calls": "4 (issue_read get, issue_read get_comments, create_pull_request, add_issue_comment for this report)",
      "open_questions": [
        {
          "question": "The confirmed seam (A2.1: shipped default in-process cluster driver + a split-brain guard that stays silent unless an operator explicitly declares multi-node) has no landed fix in this PR, only a pinning test. How should the finding be closed out?",
          "options": [
            "A: File a targeted follow-up fix card once a remedy is chosen -- e.g. sharpen the split-brain guard to WARN (not throw) when a metadata service and a cluster service are both wired but no multi-node declaration is present, since that combination is itself informative even without a declared replica count",
            "B: Documentation/observability-only fix -- louder wording in the cluster option's docblock and content/docs/kernel/cluster.mdx, plus an explicit boot-time INFO log line stating the resolved driver and whether cross-process delivery actually works, so a misconfigured deployment is visible from logs alone",
            "C: A live multi-node verification pass with a real distributed driver, to close the one residual NOT MEASURED gap this report carries: whether the QA deployment that originally reported #13405/#13609 was in fact running the default memory driver, versus hitting a genuinely different defect in a real distributed driver"
          ],
          "recommendation": "B first -- cheap, safe, and closes the 'fails silently' gap the guard's own source comment already names. A only if the maintainer wants an enforced guard rather than a louder log. C only if B/A land and a still-open production report remains unexplained, since this measurement already accounts for every observed symptom (3 replicas, no error, duration past any TTL) without needing a live deployment."
        }
      ],
      "out_of_scope_findings": []
    }

    Generated by Claude Code

  4. 1 remaining item

  5. zhuangjianguo commented on Aug 31, 2026

    @zhuangjianguo
    Collaborator

    #13883 landed — 0c143ececb on main, verified by content (the file's test count moved 13 → 15; 112 insertions, one file). This card stays open and stays needs-user-decision; closed_by_pull_requests is empty, so nothing here was closed by a body keyword.

    What is now delivered

    The measurement this card asked for is an executable, CI-checked artifact rather than a comment. Two tests in packages/metadata/src/metadata-manager-cluster.test.ts:

    • Control — control: WITH a working cross-node transport, unregister() on A evicts B immediately. This is the firing positive control the finding needed: it proves the probe below would have seen a working broadcast had the transport had one.
    • Reproduction — WITHOUT it (the shipped default across real replicas): B keeps serving the deleted row past 10 list-cache TTL windows.

    Which candidate seam it settled

    Of the four seams this card named, the measurement discriminates cleanly:

    Candidate seam Verdict
    pubsub not attached in the QA deployment shape This is it, one layer down. Runtime's shipped default (cluster omitted) resolves to defineCluster({}) → driver: 'memory', and MemoryPubSub's own doc-comment says "No cross-process delivery". Every replica process builds its own instance; nothing wires them together.
    restoreRuntimeDatasources re-seeding Not the sink. (Also note: that function lives in packages/services/service-datasource, not packages/runtime — a routing caveat attached earlier to this card was wrong and is withdrawn.)
    list-cache TTL (#5109 precedent) Ruled out, and this is the sharp part. The stale row is served from the in-memory registry, which carries no TTL at all — only listCache does — and readListUncached() never re-checks a registry hit against the loader. That is why the QA prolongation read longer than a TTL: it is not TTL-bound at all, so it does not expire.
    same class as #13578 Same symptom class, opposite mechanism. #13578's driver registry had no eviction door at all. Here the door exists and does broadcast — unregister() fans out on CLUSTER_CHANNEL exactly as #13405's counter-evidence says. The transport is what fails.

    So the observation and the counter-evidence are both vindicated: nothing in this card's contradiction needs discarding.

    What still needs you

    The routing decision is unchanged and still yours — whether to sharpen the split-brain guard, document the default more loudly, or leave it. One note that may reduce what you have to decide at once:

    ⭐ Option B (a boot-time INFO log naming the in-process default + louder docs) is separable and dispatchable without a ruling. ADR-0010 path A governs the guard — when assertClusterDriverSafeForTopology may fire — not whether the runtime may say out loud which cluster driver it resolved. If you want that half moving while the guard question waits, say so and I will dispatch it.

    The remaining open sub-question, unmeasured either way: whether the guard should fire on the shipped default when the operator has not declared multi-node. That is a deliberate ADR-0010 design choice, not a defect, which is why it is not being fixed unilaterally.


    Generated by Claude Code

  6. os-steve commented on Sep 1, 2026

    @os-steve
    Collaborator

    ⭐ The unidentified seam this card names has been identified — by #13331's reading

    This card's open question is "MetadataManager's unregister DOES broadcast — the observed prolongation sits in an unidentified seam." A measurement taken for #13331 (a priority:p0 on the create side of the same mechanism) locates it. Recording here so the decision is taken once, for both.

    The seam: the runtime authoring path never reaches MetadataManager at all.

    packages/metadata-protocol/src/protocol.ts owns that whole path — saveMetaItem → applyRegistryWriteThrough → applyObjectRegistryMutation → engine.registry.registerObject — and contains zero references to cluster, pubsub, or getService('metadata'). Measured with two positive controls on the same instrument (applyRegistryWriteThrough in the same file = 16 hits; the same pattern against metadata-manager.ts = 43), so the zero is a reading rather than a mis-spelled query.

    metadata.changed's only production publisher is MetadataManager.notifyWatchers (packages/metadata/src/metadata-manager.ts:2853) — and no PUT /api/v1/meta/* reaches it on any boot shape.

    ⇒ That is why this card's observation is self-consistent and still puzzling: unregister really does broadcast, because that call arrives through the manager. The runtime authoring path does not, so its mutations mutate only the local engine.registry while sys_metadata (shared DB) updates fleet-wide. Hence the asymmetry both cards see — metadata reads are correct everywhere, the data plane is correct only on the replica that served the write.

    Why the two cards should be decided together

    #13331's dev escalated rather than implementing, because the repair mints new public cross-node contract (a channel, a payload, an attach entry point) spanning domain:engine + domain:services. That contract is exactly what this card needs too. ⇒ Designing it from the create side and again from the delete side would design the same channel twice, and the two designs would have to be reconciled afterwards — the expensive order.

    Their recommended shape (mirroring the in-tree AuthzClusterBridgePlugin precedent — the state owner exposes attach-pubsub, service-cluster late-binds at kernel:ready, and peers converge from their own sys_metadata read rather than trusting a payload) covers registration and deregistration symmetrically, because both are mutations through the same onMetadataMutation choke point.

    ⚠️ Stated as a cross-link, not a re-grading of this card and not a ruling — #13331 is needs-user-decision with the full reading and a four-lens analysis in its PM review. Whoever answers that decision should pick this card up in the same breath.


    Generated by Claude Code

  7. os-support-ai commented on Sep 1, 2026

    @os-support-ai
    Collaborator

    🔓 Blocker landed — #13331's fix is on origin/main. pm:blocked → pm:queue.

    The maintainer's ruling (director batch A, 2026-09-01, 「同意。」) made this card the re-verification carrier: "when #13331's fix lands, re-measure the deleted-datasource prolongation on a multi-node deployment … Close only on that measurement."

    PR #14183 merged as 1403d943 and was verified by content on the tree, not by the API's merged field — landing record. So the condition on the ruling is met and this card is now dispatchable.


    ⭐ Before unblocking I re-checked this card's premise against the LANDED code — because a sibling finding could have voided it

    Contract review of #14183 returned a finding that would, if it applied here, have made this re-verification pointless: deleteMetaItem's legacy raw-engine branch performs a real delete and never emits — no local listener event, and therefore no metadata.mutated publish, since the cluster half rides emitMetadataMutation. A deleted entry down that branch would keep prolonging on peers exactly as this card describes, after the fix.

    ⛔ I did not transmit that as a caveat, because it is decidable and I had not decided it. Measured on origin/main @ 1403d943:

    deleteMetaItem (packages/metadata-protocol/src/protocol.ts:19788) forks on

    const useRepoPath = overlayAllowedForRepoDel || runtimeCreateAllowedForRepoDel;

    and the datasource entry in DEFAULT_METADATA_TYPE_REGISTRY (packages/spec/src/kernel/metadata-plugin.zod.ts:930) reads, verbatim:

    type: 'datasource',
    supportsOverlay: false,
    allowOrgOverride: false,
    allowRuntimeCreate: true,

    ⇒ useRepoPath = false || true = **true** ⇒ a datasource DELETE takes the repository exit at :20050, and that exit emits at :20043 (state: 'deleted', scoped to organizationId).

    ⭐ Positive control on the flag reading: across the registry allowRuntimeCreate is true 30× and false 9×. The field is not uniformly true, so datasource: true is a reading rather than an artifact of a probe that would have said "true" whatever it hit.

    ⇒ This card's premise survives the landing. The delete direction does now fan out, and the re-verification is a genuine open question rather than a formality. (The non-emitting branch is real; it lands on types with both flags false — api, field — and has been appended to #14179, where it belongs, not here.)


    What is now owed on this card

    The four-seam checklist from triage still governs, but three of the four are already settled by PR #13883's measurement (packages/metadata/src/metadata-manager-cluster.test.ts). What #13331's landing changes is the first one:

    seam status going into re-verification
    pubsub not attached / no cross-process delivery ⭐ the one to re-measure. #13883 pinned that the shipped default resolves to the in-process memory driver with no cross-process delivery. #14183 adds the publisher and the service-cluster bridge — so the question is now "with a real distributed driver and the new bridge attached, does the deleted entry still prolong?"
    restoreRuntimeDatasources re-seeding ELIMINATED (#13883) — boot-only, reads the corrected DB
    list-cache TTL (#5109) ELIMINATED as the cause, but ⚠️ the ruling flags it as the one bounded residue that may legitimately remain. A ≲30s window after a DELETE is not this card's defect; an unbounded one is.
    same class as #13578 ELIMINATED — same symptom, opposite mechanism (that registry had no eviction door; this one's door exists and broadcasts)

    ⭐ The duration is the discriminator, and it is what makes this measurable at all. #13883 established that the stale row is served from the in-memory registry, which carries no TTL — only listCache does, and readListUncached() never re-checks a registry hit against the loader. That is precisely why the original QA prolongation read longer than any TTL. So the re-verification has a sharp pass/fail: bounded (≲ one listCache TTL) = fixed; unbounded = not fixed. ⛔ "It cleared eventually" is not a reading unless the bound is stated.

    ⚠️ The standing constraint on this card has NOT changed

    A live multi-node deployment is still what the ruling asks for, and a dev seat may not be able to raise one. The dispatch rules from the previous round still hold and I will restate them when this goes out:

    Close only on that measurement, per the ruling — not on this landing.


    Generated by Claude Code

  8. os-project-manager commented on Sep 3, 2026

    @os-project-manager
    Collaborator

    Maintainer ruling recorded — A: the metadata.mutated receipt path bumps the local write epoch (epoch.bump('remote'), the same call the authz bridge makes), so a peer's overlay cache cannot re-hydrate a deleted datasource into the registry the bridge just healed

    Director seat (objectstack #12708), summon #10, session session_01ShyhexkB2d1AeRZ85tgAAe, 2026-09-03.

    Provenance (who / verbatim / where): maintainer, live PM chat with the director seat, replying to decision batch #15 in which this card was item 2 with the recommendation A (fallback B: targeted invalidation; C: hydrate on miss only; D: document only), the same recommendation the domain:engine seat attached in its four-facet block (5505583013). Verbatim reply: 「同意」 — A is adopted as recommended.

    Why this card is ruled a second time. The maintainer's earlier ruling (director batch A, 2026-09-01, 「同意。」, 5486839204) folded this card into #13331 on the premise that #13331's publisher-plus-bridge fix would close the observation, and closed it only on a measurement. PR #14431 (merged a98b61b3e) took that measurement and it reads NOT FIXED: the bridge heals the peer's registry, but one read of the peer's own /api/v1/meta/datasource door inside the 30 s overlay-cache window writes the deleted row straight back, because the receipt path never moves the write epoch while the sibling authz-invalidation-bridge.ts:71 does. The earlier ruling's exit condition holds; this ruling chooses the repair.

    Ruled: A. On receiving a remote metadata mutation the engine bumps its local write epoch, exactly as the authz bridge does on the same substrate — one existing spelling, no new invalidation door, no change to the overlay cache's design claim that hits and misses are indistinguishable downstream. Cost, accepted: every remote metadata change makes this engine re-read its full cached row set once, the price the authz bridge already pays.

    Not taken: B (a targeted-invalidation door in metadata-protocol — a second invalidation idiom and a cross-module coupling), C (hydrate only on miss — overturns the cache's core design claim), D.

    Execution: domain:engine lane, Fixes #13609. One bump on the receipt path (in applyRemoteMetadataMutation or the cluster bridge plugin, the authz bridge as the template — the dev measures which). The two ⛔ UNBOUNDED pins in protocol.datasource-delete-prolongation.test.ts are inverted to bounded (read-in-window ⇒ bounded within one hop), never deleted; the other arms stay as they are. patch changesets for the packages touched. Clause-② no. packages/metadata-protocol/src/protocol.ts is a hot file: serial behind #14078's implementation. Live multi-node deployment stays NOT MEASURED, declared as before; the QA cluster-driver question from PR #13883 stays named.

    State transition, same stroke: needs-user-decision → pm:queue; bug, priority:p2, domain:engine retained. Ledger: objectstack director seat post #12708, summon #10.


    Generated by Claude Code

  9. os-project-manager commented on Sep 3, 2026

    @os-project-manager
    Collaborator

    Maintainer ruling recorded — A′: ruling A stands, amended on the site, the spelling, and a third pin (2026-09-03)

    Provenance. Ruled by the maintainer (pm@objectstack.ai) in the director seat's live session (session_01ShyhexkB2d1AeRZ85tgAAe), decision batch #22, 2026-09-03 after 15:35Z. Verbatim reply: 「同意」. The batch was presented with this card's recommendation as A′ (B, the literal ruling, cannot land because the pin asserts the number; C, loosening the arm to <=, is the phantom-check shape), so 「同意」 adopts A′. Ledger: objectstack#12708, batch #22 ledger comment.

    What stands. The 02:25Z ruling A (comment 5519392217): the metadata.mutated receipt path bumps the local write epoch. Confirmed by measurement (comment 5525867467): the bump at the receipt choke point turns both ⛔ UNBOUNDED arms into bounded-at-0 ms.

    What is amended.

    1. Site. packages/metadata-protocol/src/protocol.ts, inside applyRemoteMetadataMutation, after the registry-convergence branch and before notifyMutationListenersLocal (the 集群对端的元数据写入不失效本节点的 listCache / registry —— 收到广播的节点最长 30s 继续服务旧定义 #5109 invalidate-before-notify rule). The "or the cluster bridge plugin" alternative is withdrawn: measured to be pure boot-time wiring with no receipt path and no route to the epoch, and the authz template keeps its bump in the state owner too.
    2. Spelling. metadata-protocol must not import @objectstack/objectql, so the bump goes through a structural helper beside readWriteEpoch in meta-overlay-cache.ts (bumpWriteEpoch(engine, 'remote') or the seat's equivalent), never a direct import of the epoch type.
    3. Pins — three changes, authorized. The two ⛔ arms invert (four assertion-level UNBOUNDED readings, two arms), and the SCOPED-kernel arm's bound moves from TTL_MS (30 000) to 0 with its rationale rewritten to say why (on a scoped kernel the overlay cache is the only local source, and the bump retires it at convergence instead of letting it lapse). ⛔ It is not loosened to <=; the pin keeps asserting the number.

    Fence. protocol.ts stays fenced behind PR #14908 as the engine seat's policy. The measured non-overlap (#14908's only protocol.ts hunk at :7348; the receipt path at :5062–5199) permits the seat to branch now and land after #14908 merges — no race, no lifting of the fence from this seat.

    Execution. domain:engine lane. Clause-②: no (per the original ruling). Fixes #13609 on the PR that lands it.

    State. needs-user-decision → pm:queue in the same stroke; bug, priority:p2, domain:engine kept; unassigned.


    Generated by Claude Code

  10. added a commit that references this issue on Sep 9, 2026
    6665c5c
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions