Skip to content

008G: mission-local repair reachability proof - #75

Merged
Thunderkill016 merged 9 commits into
mainfrom
devin/m008g-repair-reachability
Oct 1, 2026
Merged

Thunderkill016 merged 9 commits into
mainfrom
devin/m008g-repair-reachability

Conversation

@Thunderkill016

@Thunderkill016 Thunderkill016 commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Summary

Mission 008G replaces the full-task-registry repair-channel check in
deriveCorrectionEpisodes with an explicit, machine-readable
mission-local repair reachability proof (control-room follow-up to
008F, PR #74).

  • task-resolver.js (new): mission-scoped servable/pickTask/
    optionsFor + verified-event consumption extracted from the candidate
    generator; excludeTaskIds enables counterfactual reachability. B0
    never passes exclusions — routing is byte-identical.
  • correction-episodes.js: evidence-only; retestReservedTaskIds and
    the registry scan removed.
  • repair-proof.js (new): deriveMissionRepairPlan(s) classifies EVERY
    live CORRECTION/REFRESH route on its natural next serve (usable / inert
    / dangerous); complete iff ≥1 usable (mission-local, clean,
    repair-eligible, non-probe, covers ALL remaining fns) and no dangerous
    route (R1 + R2 review); per-fn witness streams kept for audit; deriveRetestReservations turns
    complete proofs into the withheld set. Reservation table:
    OPEN never · REPAIRING/RELAPSED iff complete proof · REPAIRED_WAITING
    direct · RETEST_DUE/VERIFIED never.
  • policies.js: hardFilter takes the reservation set; policyB1 wires
    the proof; decisions carry repairPlans for audit.
  • validator.js: B1 decisions independently re-derive proof +
    reservations (memoized rebuild) — forged serves on reserved probes fail
    closed.
  • package.json: wires vnext-next-for-you-runtime.test.mjs into
    npm test — the suite was orphaned from verify/CI since 008E.

Evidence

  • Runtime suite: 342 checks PASS (+39 new: 3 positive proofs, state
    table, cross-mission/ghost-registry/wrong-cap/wrong-fn/stale-rev/
    repair-bound/failure-ceiling/repick/multi-fn/relapse attacks,
    MRP-SERVEDNEXT split-coverage pair (R1), MRP-XROUTE cross-route
    conflict pair (R2), registry-invariance metamorphic, validator forgery).
  • Differential: B0 corpus 1792 rows {EXPECTED:1009, MATCH:783}
    byte-identical; B0↔B1 MATCH:1744, CORRECTION_RETEST_SURFACE_RESERVED: 42, CORRECTION_RETEST_DUE:6, violations 0, BUG 0.
  • Perf @2k: b1 124.1ms vs b0 125.9ms; proof reuses gen.resolver —
    no second candidate pass or model replay.
  • swe:verify PASS @ 5b4f0ae and 3870b9f; direct verify:full green @ fc50e9d (R1) and on the R2 tree committed as 0bbfa0b (factory refuses re-verify on a DONE mission).

Test plan

  • npm run verify:full green on ending SHA
  • differential corpus + B0↔B1 classes clean
  • perf profiled
  • CI green on 9d3690c and R1 fc50e9d
  • CI green on R2 head 0bbfa0b (push 36831383522 + pr 36831387521)
  • control-room policy review

Not merged autonomously — awaiting review; user is merge authority.
B1 remains shadow/experiment; production route serves B0.

Generated with Devin

Thunderkill016 and others added 6 commits October 1, 2026 12:40
Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Replaces the full-registry repair-channel check with a per-episode,
per-function proof through the shared task resolver.

- task-resolver.js: extracted mission-scoped servable/pickTask/
  optionsFor machinery + verified-event consumption; adds
  excludeTaskIds for counterfactual reachability. B0 never passes
  exclusions — byte-identical routing.
- correction-episodes.js: evidence-only; retestReservedTaskIds (the
  registry-scan false positive) removed.
- repair-proof.js: deriveMissionRepairPlan(s) proves ∀ missing fn ∃
  mission-local, hard-filter-clean repair witness under probe
  exclusion; deriveRetestReservations turns complete proofs into the
  withheld probe set.
- policies.js: hardFilter takes the reservation set; policyB1 wires
  proof between episode derivation and applyFilters; decision carries
  repairPlans for audit.
- validator.js: B1 decisions independently re-derive proof +
  reservations (memoized rebuild); forged serve on a reserved probe
  fails closed with correction_retest_surface_reserved.
- tests: +32 008G checks — clock/price/direction positive proofs,
  state table, cross-mission/ghost-registry/wrong-cap/wrong-fn/
  stale-rev/repair-bound/failure-ceiling/repick/multi-fn/relapse
  attacks, registry-invariance metamorphic, validator forgery.
- package.json: wires vnext-next-for-you-runtime into npm test —
  the suite was orphaned since 008E and never ran under verify/CI.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

Copy link
Copy Markdown
Owner Author

Policy review of PR #75 head 9d3690c64c67d57e6311c772981335031e268002.

Exact-head CI is green (push 36825803861, PR 36825837191, both SUCCESS), base main is still e98a6e7a8e9017cb83fae4600fd160ef36084eb5, and the main architectural move is correct: correction episodes are evidence-only, mission-local routing owns reservation, the shared resolver removes registry-vs-reachability drift, and B0 parity remains intact.

DO NOT MERGE #75 YET. I found one reachability blocker in the proof contract, plus durable-truth cleanup.

BLOCKER — complete proves “a clean task exists in the route stream”, not “the route can actually reach it under current policy state”

In deriveMissionRepairPlan(), each function’s witness stream is enumerated with:

  • mission-local tasks;
  • probe exclusions;
  • per-task hardFilter.

But final completeness accepts any witness where:

missionMember && coversFunction && hardFilterClean && !consumesFreshRetestSurface

It does not require servedNext.

That makes the proof stronger than what the live route actually establishes.

Concrete counterexample:

  • episode misses functions F1 + F2;
  • mission has remediation A covering F1, then remediation B covering F2;
  • both are in the route stream;
  • decisionContext is already at repair_bound - 1;
  • under the current state, both A and B individually look hardFilterClean because each hypothetical is checked against the SAME pre-action context;
  • the actual route can serve only A next;
  • after A is consumed, the repair action count hits the bound;
  • B is no longer reachable.

Current proof can still report both per-function witnesses clean, complete=true, and reserve the retest probes.

So the implementation proves independent static existence, not a jointly feasible policy path.

This is exactly the distinction 008G exists to enforce: task existence / stream membership != reachable repair channel.

Required narrow fix

Do not build a general state-space planner in 008G.

Use a conservative v1 proof:

For REPAIRING / RELAPSED, reservation is allowed only when the actual next repair serve under probe exclusion is validator-clean and that served-next surface covers ALL currently remaining functions.

In practice:

  • servedNext === true must be load-bearing, not audit-only;
  • the same actual next surface may witness multiple functions;
  • every remaining function must be covered by a clean served-next repair surface;
  • if separate future repair actions would be required, return mission_repair_channel_unproven for v1 rather than pretending the sequence is reachable.

This is conservative but correct, and the three current real paths should still pass because their authored remediation surfaces already cover the attributed function set.

If you prefer to support multi-step repair later, that needs an explicit bounded transition simulation that updates:

  • task consumption,
  • actionsChosen / repair bound,
  • observed fail streak,
  • episode state,
    after every simulated serve. That is out of scope for this patch.

Required regression

Add an adversarial fixture:

  1. one episode missing F1 + F2;
  2. remediation A covers F1;
  3. remediation B covers F2;
  4. mission order serves A before B;
  5. decisionContext has exactly one repair action left before the bound.

Expected:

  • A may be a clean served-next witness;
  • B must NOT count as currently reachable;
  • complete=false;
  • reservations empty.

Also add the positive dual:

  • one served-next remediation task covers F1 + F2;
  • proof remains complete and reservations apply.

This test should fail on the current head.

IMPORTANT FOLLOW-UP — partial repair semantics in REPAIRED_WAITING

Not a merge blocker for 008G if documented, because it is inherited from 008F and matches the mission’s current state table, but record it explicitly:

deriveCorrectionEpisodes enters REPAIRED_WAITING after an independent success covering ≥1 missing function, not necessarily all of them.

008G then directly reserves in REPAIRED_WAITING without a repair-channel proof.

Therefore the new ∀-function proof is currently strongest only in REPAIRING/RELAPSED; it does NOT yet prove that every missing function was actually repaired before the waiting clock starts.

Before B1 is generalized to broader multi-function episodes, we need either:

  • replay-derived repairedFunctions / repair-demonstration provenance and wait only after full immediate repair coverage; or
  • a separate explicit policy decision that partial repair before delayed verification is intended.

Do not silently claim 008G solves that older semantic question.

Durable truth cleanup required after the code patch

Current committed REPORT still has:

  • all acceptance boxes unchecked despite DONE;
  • CI described as pending;
  • ending SHA 3870b9f while PR head is 9d3690c.

PR body likewise still shows CI unchecked even though exact-head CI is green.

After the final code patch:

  • update REPORT acceptance state honestly;
  • distinguish code-final SHA vs report/docs head;
  • record exact-head CI run ids;
  • update PR body.

Accepted

I am satisfied with:

  • deriveCorrectionEpisodes no longer owning reservations;
  • mission-local task resolution through the shared resolver;
  • probe exclusion / alternate repick mechanism;
  • cross-mission and registry-invariance attacks;
  • wrong-cap/wrong-function/revision/bound/ceiling attacks;
  • validator independently re-deriving reservations;
  • runtime suite now being wired into npm test;
  • B0 corpus parity and unchanged B1 divergence counts;
  • performance profile and no second learner-model/generator pass.

Return:

008G REACHABILITY FINAL PATCH RESULT
HEAD
SERVED-NEXT PROOF
JOINT-FEASIBILITY ATTACK
MULTI-FUNCTION POSITIVE
CLOCK/PRICE/DIRECTION
B0 PARITY
B1 DIFFERENTIAL
PERFORMANCE
DURABLE REPORT
VERIFY:FULL
CI
OPEN LIMITATIONS

End:
PR #75 NOT MERGED — AWAITING FINAL CHATGPT CLEARANCE

Do not start the next mission.

Thunderkill016 and others added 2 commits October 1, 2026 13:58
Review found the per-function witness aggregation could claim false
reachability: F1→A and F2→B each get a stream witness, but the route
serves ONE task next and serving A may exhaust the repair bound — B is
never actually reachable.

Narrow patch: REPAIRING/RELAPSED completeness now requires SOME live
repair route whose actual next serve under probe exclusion is
mission-local, hard-filter-clean, non-probe, and covers EVERY remaining
function. Per-function witness streams remain as audit detail only.

- plan.servedNext[] records each route's real next serve under exclusion
  with coversAllRemaining
- MRP-SERVEDNEXT regression pair: split coverage under bound pressure →
  complete=false, reservations=[]; single served task covering all
  missing fns → proven, probes withheld
- differential unchanged: 1792 {1009,783}; B1 1744+42+6, 0 violations —
  real-path remediations cover their full missing-fn sets
- REPORT: acceptance boxes + limitation note (REPAIRED_WAITING enters on
  ≥1-fn repair; multi-function generalisation deferred)

Runtime suite 338 checks.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

Copy link
Copy Markdown
Owner Author

R1 re-review of PR #75 current head e259be864c7e98e6b5870eea1406599ec479bafe.

Exact-current-head Verify FlashDay #310 is now SUCCESS. The served-next patch fixes the specific split-witness bug from the previous review, and the new negative/positive pair is good.

DO NOT MERGE YET. One reachability hole remains.

BLOCKER — complete now proves “some route has a safe servedNext”, but B1 may actually choose a different repair route

The R1 code computes one servedNext per live route:

  • CORRECTION
  • REFRESH

and then does:

plan.complete = plan.servedNext.some(...coversAllRemaining...)

That is still not the actual policy-next proof.

In B1, CORRECTION and REFRESH are both REPAIR tier candidates. pickOrdinal ranks:

verified_failure_on_demonstrated (REFRESH)

above:

open_attributed_gap (CORRECTION)

So when both are live, REFRESH normally beats CORRECTION.

A counterexample remains possible:

  1. episode missing F1 + F2;
  2. burned retrieval B is first in mission order and covers only F1;
  3. remediation A is later and covers F1 + F2;
  4. with fresh probes excluded:
    • REFRESH.servedNext = B (clean, partial);
    • CORRECTION.servedNext = A (clean, covers all);
  5. some(...) sees A and marks complete=true;
  6. actual B1 ordinal policy serves REFRESH/B first;
  7. if that consumes the last repair action, A is never reachable.

That is the same class of error one level higher:
route existence != policy-reachable next action.

There is an even simpler variant: REFRESH can repick an unburned retrieval surface that is not a retest probe because it does not cover the missing function. Current proof drops it entirely at:

if (!t || !repairEligible(t, burned)) continue;

then CORRECTION can still prove complete=true, even though REFRESH may be the real policy winner and spend a repair action without repairing the episode.

Narrow v1 fix — no state-space planner

Stay conservative.

For REPAIRING / RELAPSED:

  1. record the actual served-next verdict for every live repair route, including a route whose next task is not repair-eligible;
  2. a route is safeNext only when its actual next task is:
    • mission-local;
    • hard-filter-clean;
    • repair-eligible;
    • non-probe;
    • covers ALL remaining functions;
  3. reservation may be proven only if:
    • at least one live route has a clean served-next; AND
    • every clean live repair route that could be selected at this state has safeNext=true.

For this conservative v1, you do not need to reproduce pickOrdinal inside repair-proof. Requiring all currently clean live CORRECTION/REFRESH routes to be safe is stronger than necessary but sound. The current clock/price/direction paths should continue to pass because both live routes converge on the same full-coverage remediation after probe exclusion.

Do not ignore/skip a live route merely because its next task is not repairEligible; record it as an unsafe route and let it falsify the proof if hard-filter-clean.

Required regression

Add a route-conflict attack:

  • F1 + F2 missing;
  • burned retrieval B covers F1 only and is earlier for REFRESH;
  • remediation A covers F1+F2 and is CORRECTION's served-next;
  • both routes are hard-filter-clean;
  • context has one repair action left.

Assert:

  • REFRESH servedNext = B, safeNext=false;
  • CORRECTION servedNext = A, safeNext=true;
  • current ordinal ordering would prefer REFRESH (pin this fact or at least its preference relation);
  • complete=false;
  • reservations empty.

Positive dual:

  • both live routes' served-next resolve to a full-coverage repair surface;
  • complete=true.

This should fail on current R1 because some(...) accepts the safe CORRECTION route.

Durable truth

After the code patch, update REPORT again:

  • runtime count is now 338, not the still-recorded 335;
  • current docs head is e259be8;
  • exact-current-head CI #310 is SUCCESS;
  • distinguish code-final fc50e9d from docs/report head e259be8.

PR body also still shows the old 335 checks / old CI checkbox state.

Accepted from R1

  • split F1→A / F2→B is now correctly rejected;
  • single served-next covering all functions is correctly accepted;
  • per-function streams are now audit-only;
  • the inherited REPAIRED_WAITING partial-repair limitation is explicitly documented;
  • B0/B1 differential counts stayed stable;
  • runtime suite is in the verify chain;
  • exact-current-head CI is green.

Return:

008G ROUTE-SELECTION FINAL PATCH RESULT
HEAD
ALL-LIVE-ROUTE SAFETY
ROUTE-CONFLICT ATTACK
POSITIVE DUAL
CLOCK/PRICE/DIRECTION
B0 PARITY
B1 DIFFERENTIAL
DURABLE REPORT
VERIFY:FULL
CI
OPEN LIMITATIONS

End:
PR #75 NOT MERGED — AWAITING FINAL CHATGPT CLEARANCE

Do not start the next mission.

R1 accepted the proof when one route's served-next covered all missing
functions, but the policy may prefer a different route whose next serve
is partial and burns the last repair action. Each live CORRECTION/REFRESH
route is now classified on its natural pick (usable / inert / dangerous);
complete requires a usable route and no dangerous one. Adds MRP-XROUTE.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

Copy link
Copy Markdown
Owner Author

Final policy clearance on PR #75 head 0bbfa0b24298b3d0fb3fc8a17277fc39b2e8e779.

Independent final review confirms the R2 route-selection patch closes the remaining reachability hole:

  • every live CORRECTION/REFRESH route is evaluated on its actual natural next pick;
  • a live clean route that is non-repair-eligible or only partially covers the remaining function set is classified dangerous and falsifies the proof;
  • reservation now requires ≥1 usable route and zero dangerous routes;
  • the MRP-XROUTE conflict/positive-dual regressions pin the cross-route case;
  • mission-local / registry-invariance / served-next / validator-forgery attacks remain green;
  • B0 frozen corpus remains byte-identical and B1 differential counts are unchanged;
  • runtime suite is now in the normal verify/CI chain;
  • exact-current-head Verify FlashDay #312 is SUCCESS;
  • base main remains e98a6e7a8e9017cb83fae4600fd160ef36084eb5, with no conflicting semantic changes.

The inherited REPAIRED_WAITING partial-repair limitation is correctly documented and is the next semantic target; 008G does not overclaim to solve it.

008G is cleared for merge. No B1 product rollout, public deploy, or Firestore-rules deployment is authorized by this merge.

@Thunderkill016
Thunderkill016 merged commit 85ce410 into main Oct 1, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant