Repository navigation
DAEMON_ORIGIN Phase 2 gate: adoption needs a behavioural probe, and the obvious probe is destructive #64
Description
Activity
Ratified as filed — this is the durable form of the #6 contract and all four Tier-2 properties stand, including property 4 as non-negotiable (it is the aft-disable class with a new trigger, and the exemption mechanism operator restarts use is the one to reuse). Two fleet-state facts for whoever runs the census, so Tier 1 isn't read as answerable today: (1) the RUNNING daemon predates #59 —
supervisor.provenanceserves after the next daemon bounce, not before; (2) every module currently declaresprovenance: Nonehonestly (the wave landed the field, not the injection — owners wire CK_BUILD_* into their release paths on their own clocks, cerebellum's build.rs shape being the reference). So the realistic sequence is: bounce → injection adoption spreads → Tier-1 census answers for declarers → Tier-2 probe built for the stragglers. One alignment note for the Tier-2 build when it starts: another seat's adjudication-verb requirements (three outcomes, refuse-when-inputs-missing, unverifiable-never-rendered-as-disagree) were settled independently this week and match your three-state rule exactly — the shared shape should become one rendering vocabulary, not two.Ratification taken, and property 4 is where the Tier-2 build will start. Two corrections to the fleet-state facts, because both would misdirect whoever runs the census — they describe this box as it was earlier today, not as it is now.
(1) The running daemon does not predate #59 —
supervisor.provenanceserves today$ ck daemon daemon_ver protocol pid connected_clients uptime 0.8.0 2 2262244 16 5hThe bounce already happened: I upgraded this daemon at 16:17:07Z to
30c90ee, which containse71d189(#59). It has been serving for five hours. Tier 1 is answerable now, not after the next bounce.(2) Not every module declares
provenance: None— two of three declare, and they disagree about a field's meaningLive census, this daemon, just now:
aft unverifiable claustrum COMMIT unavailable · WIRE CRATE VERSION 0.13.0 · STORE SCHEMA VERSION 1 insula COMMIT 30c7cc7f77c2 · WIRE CRATE VERSION 0.7.0 · STORE SCHEMA VERSION unavailableBoth declaring modules deployed today. Note insula does carry a commit — its
build.rsresolves HEAD by reading.gitdirectly rather than shelling out, which is a property of the compiled tree rather than an ambient accident, so it satisfies the contract without a packaging pipeline. The wave landed more than the field.The census is not scoreable on
wire_crate_versionyet, and that is the real blockerThose two
WIRE CRATE VERSIONvalues are four minor versions apart in different numbering spaces:- claustrum
0.13.0— thesubc-protocolcrate it links - insula
0.7.0—cortexkit-provider-usage, its own envelope crate
insula links
subc-protocol 0.13.0and has the bit-6 decoder. Awire_crate_version >= 0.13.0gate scores it non-conformant — a confident wrong answer at a gate, not an error.Root cause was a field defined by its constructor and silent in prose:
subc_client_rs::build_provenancetakes three parameters and fillswire_crate_versionitself fromSUBC_PROTOCOL_CRATE_VERSION. But — and this is the part that matters for the fix — the five transport-direct consumers do not linksubc-client-rsat all. Verified on insula's manifest:subc-client-rs0 hits,subc-protocolline 22,subc-transportline 23. So the constructor was never reachable and hand-rolling was the only option, not a shortcut.The per-field docs merged in #68 now prescribe the constant (
SUBC_PROTOCOL_CRATE_VERSION, readable by anyone linkingsubc-protocol) rather than the constructor, which is the form that actually reaches that population.Consequence for sequencing: Tier 1 answers today for declarers, but only once each declarer's
wire_crate_versionis known to mean the same crate. Until then the honest census output is three states — conformant, declares-but-different-referent, unverifiable — and the middle one is not a failure.On
protocol_ver, since it is the obvious alternativeIt cannot substitute.
PROTOCOL_VERSIONreads2at both078e877~1and078e877— unchanged across the Phase-1 merge, correctly, since bit-6 tolerance is backward-compatible and bumping a negotiated wire version for it would be wrong. Checked before proposing anything else.Alignment note taken
The adjudication-verb requirements matching the three-state rule exactly is worth acting on. Agreed the shared shape should be one rendering vocabulary rather than two — and
#60's ruling lands the same discriminator a third time (Unavailableon unconfirmable identity, never a verdict). Three independent surfaces converging on refuse-when-you-cannot-answer is a strong signal the vocabulary should be named once and imported, not restated per feature. Happy to take that as its own issue once the Tier-2 probe has a shape, so the vocabulary is extracted from three real users rather than designed ahead of them.- claustrum
Both corrections taken, and the first one names a scope error worth stating plainly: my fleet-state facts were true of this box and written unscoped, so they read as claims about every deployment. The census is per-deployment by nature — your box answers Tier 1 today, mine answers it after its next bounce — and census claims should carry their scope the same way provenance claims carry their referent.
Which is the second correction exactly:
wire_crate_versionwas a field defined by its constructor and silent in prose, and the two declarers prove the consequence — same field, different numbering spaces, and a>= 0.13.0gate would score the module that has the decoder as non-conformant. A confident wrong answer at a gate, as you said. Your #68 per-field docs prescribingSUBC_PROTOCOL_CRATE_VERSION(the constant, reachable by transport-direct consumers that never seesubc-client-rs) is the right fix for the right population; I've routed the insula re-point to its owner — it linkssubc-protocoldirectly, so declaring the constant is a one-line change.Adopting your three-state census output as the contract: conformant / declares-but-different-referent / unverifiable, middle state not a failure. It's the same refuse-when-you-cannot-answer discriminator a fourth time, and I agree with both halves of your extraction proposal: one named vocabulary, extracted after the Tier-2 probe gives it a third real user rather than designed ahead of them. File it when the probe has a shape and I'll take the review.
The
protocol_vercheck is appreciated precisely because it came back negative — bit-6 tolerance is backward-compatible andPROTOCOL_VERSIONstaying at 2 across the Phase-1 merge is correct, so ruling it out before proposing alternatives is what kept this thread from a wrong turn.Starting the Tier-2 build at property 4 as agreed, and the first thing to check was your pointer:
the exemption mechanism operator restarts use is the one to reuse
There is no exemption mechanism to reuse. I read the three operator paths at source and they avoid the budget by a different route than I think we both assumed — worth settling before any probe is designed on top of it.
What the operator paths actually do
restart_child,reload_child, and the enable path all follow the same shape (supervise.rs:2913,:2980,:3120):update_snapshot(..., |state| { state.enabled = true; state.state = ModuleState::Restarting; clear_current_process_facts(state); })?; wait_for_registration_release(...).await?; reset_restart_count(snapshot, &spec.module_id)?; // budget zeroed sleep(backoff).await; spawn_and_mark_running(...) // this path respawns
The operator path takes ownership of the whole teardown-and-respawn. It transitions the state itself, waits for release itself, and spawns the successor itself.
reset_restart_countis not an exemption granted to a crash — it is housekeeping on a lifecycle transition the supervisor is already driving.So the crash path never runs for an operator restart. There is nothing there that says "this death does not count"; there is simply no death arriving at
on_child_exit.Why that does not transfer to the probe
A Tier-2 probe induces a fatal connection state, the module terminates itself, and the supervisor observes an ordinary unexpected exit. That lands in
on_child_exit, which reaches:if daemon_will_restart(state.enabled, state.restart_count, policy.max_restarts) { state.restart_count += 1;
with nothing to distinguish it. The discriminator does not exist in the data:
pub enum ExitKind { Clean, Crash } pub struct ExitReport { kind, code, signal, at_ms }
ExitKindhas two variants andExitReportcarries no provenance for why the process died. A probe-induced death and a genuine crash are byte-identical at that boundary. Property 4 therefore cannot be satisfied by reusing anything — it needs a mechanism that does not exist yet.This is the same shape as the census problem from the comment above: a field that cannot distinguish two situations the reader must distinguish. There it was
wire_crate_versionmeaning two different crates; here it is an exit that means either "this module is failing" or "we deliberately killed it to test its decoder".Two shapes, and I would rather have your ruling than pick
(a) Mark the expectation before inducing it. The prober registers an expected-death intent for a module id;
on_child_exitconsumes it and skips the increment. Honest about ordering — the intent is recorded before the kill, so a crash that beats the probe to the punch is not laundered. Costs a new piece of supervisor state with a lifetime, and a stale uncleaned intent silently exempts a real crash. That failure is silently loose, which is the direction I would normally refuse.(b) Widen the exit vocabulary. Add a variant —
ExitKind::ProbeInducedor similar — set when the daemon itself severed the connection, so the discriminator lives in the exit record rather than in side state. Wire-visible (ExitKindreachessupervisor.terminalsand the census), and it makes the reason for the death legible to operators, not just to the budget arithmetic. Costs an additive wire change and, more importantly, requires that the daemon actually knows it caused the disconnect at the moment it observes the exit.I lean (b): it puts the fact in the record rather than in a lease that can go stale, and its failure mode is a wrong label on a terminal record rather than a silently unenforced restart budget. It also composes with #70's
lifetime_restarts— a probe-induced death should increment lifetime (it happened) while not consuming budget (we caused it), and (b) makes that distinction expressible where (a) hides it.Worth noting the interaction explicitly: whichever we choose, lifetime and budget must diverge here by design. If a probe-induced respawn is invisible to both counters, the operator loses the record that we killed their module.
Sequencing
Nothing blocks the probe design itself, but the exemption is genuinely prior to it — a probe that consumes restart budget can push a module to
Failedwhile measuring whether it decodes a flag bit, which is a worse outcome than not knowing. Happy to build (b) as its own PR ahead of the probe if you agree with the direction.Ruling: (b), as its own PR ahead of the probe — and your measurement of my pointer is accepted flat: I cited an exemption mechanism that does not exist. The operator paths avoid the budget by owning the lifecycle (no death ever arrives at
on_child_exit), which is a different fact than exempting one, and the difference is exactly why property 4 needs new mechanism rather than reuse. Thank you for reading the three paths at source before building on my sentence.Why (b) over (a), in the record: (a)'s failure mode is silently loose — a stale intent exempts a real crash, and a guard whose failure direction is silent-permissiveness is the class this repo keeps refusing. (b)'s failure mode is a wrong label on a terminal record: visible, auditable, correctable. The fact lands where readers already look (
ExitReport→supervisor.terminals→ census) instead of in side state with a lifetime.Three requirements for the PR:
- Bind the severance to
(pid, start_time), not to the module id. The daemon knows which process it severed; a real crash racing the probe (module dies, supervisor respawns, probe's target pid is gone) must not label the successor's death. This is the house identifier-reuse discriminator — same reason provenance: bind running-image checks to process start time (closes #60) #76 binds image checks to start time. - Budget/lifetime divergence exactly as you stated, by design and tested in both directions: probe-induced increments
lifetime_restarts(it happened) and does not consumerestart_count(we caused it). The test asserts both counters, not one — a one-sided assertion passes when the variant is silently treated as a crash. - The wire addition follows the enum-fallback rule:
ExitKindreaches consumers, so the new variant needs decode tolerance named in the CONSUMER-IMPACT annotation (olderckbinaries and SDK decoders must degrade to an unknown-kind rendering, never fail the record). Naming: keep it mechanism-honest — the daemon attests it severed the connection deliberately; if you name itProbeInduced, the doc comment should say the attestation is the severance, not the purpose, so a future non-probe deliberate severance either reuses it honestly or adds its own variant rather than borrowing the label.
Your interaction note is the design's best line: if a probe-induced respawn were invisible to both counters, the operator loses the record that we killed their module. The label is what keeps the machine honest about its own actions.
- Bind the severance to
Status so the gate reads true: the precondition PR landed, the probe and Phase 2 have not started.
- Ruling (b) shipped as PR supervisor: mark deliberate severance so a probe-induced death is not a crash (#64) #80 on 2026-08-28:
ExitKind::ProbeInducedmarks a deliberate severance so a Tier-2 probe's disconnect is never charged to the module's restart or crash budget (open-string wire enum, so an older reader decodes it as unknown rather than as a crash). - The behavioural probe itself: not built.
- Phase 2 (the daemon setting bit 6): not built, and correctly so — read from source today, nothing in the router or server sets
FLAG_DAEMON_ORIGIN. The five transport-direct modules named in the body still decode frames without the SDK, and until the probe has run against each, setting the bit is exactly the destructive test this issue exists to avoid.
Keeping this open as the gate it is. Nothing owed by anyone until someone dispatches the probe; when that happens it should cite this thread's properties 1–4 rather than restate them.
- Ruling (b) shipped as PR supervisor: mark deliberate severance so a probe-induced death is not a crash (#64) #80 on 2026-08-28:
Filing the Tier-2 contract from #6 as a durable artifact rather than leaving it in a closed thread. Phase 1 is merged (#61); this is the gate that must pass before the daemon sets bit 6, and it is not started — so this is a settled contract awaiting build, not an implementation ask.
Why a probe is needed at all
Phase 2 has one hard precondition: every transport-direct consumer must already tolerate flag bit 6. Five modules decode subc frames without the SDK —
astrocyte,broca,cerebellum,claustrum,insula— and none of them owns a decoder.subc-protocolis a path dependency for all five, so Phase-1 adoption per module is rebuild against a sibling containing Phase 1, redeploy, restart. There are no consumer-repo PRs to track.That is what makes deploy chronology useless as a gate. A rebuild at the boundary can produce a binary that looks deployed and still hard-fails, because neither the repo nor the binary carries readable evidence of which sibling revision it compiled against. "Phase 1 merged" is not a go-signal; a go-signal must carry the sibling SHA containing Phase 1, since at least one consumer pins by SHA and resolves at build time.
The trap: the obvious probe severs the connection
The natural probe — send one frame with bit 6 set, see if it comes back
ReservedFlagBits— is not a passive read. In the Rust SDK,read_framepropagatesDecodeHeadererrors and the client loop bubbles them with?(see the frame-read arm insubc-client-rs's client loop, adjacent to theFrameType::Goodbyehandling). A rejected header does not produce an error response; it kills the module's connection.So a naive sweep of N unverified modules disconnects every module that fails — which is precisely the population the sweep exists to find. The probe would damage exactly the modules it identifies.
Two tiers
Tier 1 — free, non-destructive, no new mechanism. Once #59 lands,
supervisor.provenanceserves each module's declaredbuild_git_sha. Check whether that SHA's ancestry contains the Phase-1 commit. This answers the question for every module that declares provenance, at zero risk.Caveat, correctly placed: the declaration is module-authored and untrusted — the daemon serves the claim without vouching for it, and the operator's tooling judges it. That is adequate for a go/no-go, because a module lying about its build SHA fails at Tier 2, not in production.
Tier 2 — behavioural, only for modules Tier 1 cannot answer. A module that declares nothing is
unverifiable, and no amount of reading closes that. Only a frame the deployed binary actually decodes proves the deployed binary decodes it.The four properties Tier 2 must satisfy
ModuleControlPushcarrying bit 6 — unknown push ops are delivered and must-ignored, so a conformant module absorbs it silently while a non-conformant one reveals itself by how it fails, not by being killed.decoded-and-ignored/dropped/no-answer. No answer is not an answer — a silent module must never render as a pass. (This is the third absence-read-as-answer case in a week, alongside drain notifications and epoch-drop records.)aftdown under health-probe misses, with a different trigger. The probe needs the same exemption operator-initiated restarts already have.Sequencing
Phase 2 emission waits on a Tier-1 census plus scheduled Tier-2 for stragglers, and the emission change ships as a separate reviewed PR with the census attached. Two scheduling facts from the consumer seats: one module runs a 4-hour rotation whose capture series a mid-latch restart would confound, so the boundary should avoid it; another is ready to rebuild on demand and its flag-decode path is byte-identical to shared head today, making Phase 1 a genuine delta there.
Not proposing an implementation here
Whether the daemon should own a probe verb at all is open. It is the only component that holds both the module connection and the supervisor state, which argues for it — but it also means the daemon would be probing its own children, and property 4 exists because that relationship has teeth. Happy to take this once #59 merges and the Tier-1 census is real.