Skip to content

vnext: integrate learning-system stack through R7 - #64

Merged
Thunderkill016 merged 34 commits into
mainfrom
integration/vnext-stack-20260930
Sep 30, 2026
Merged

Thunderkill016 merged 34 commits into
mainfrom
integration/vnext-stack-20260930

Conversation

@Thunderkill016

Copy link
Copy Markdown
Owner

vNext stack integration — single candidate to main

This PR carries current main + the complete vNext development stack + the SWE Work Factory on one verified branch. Integration mission ran through the Work Factory itself — missions/003-vnext-stack-integration/ holds the mission definition, checkpoint, verification log, and REPORT.md (mission DONE).

Ancestry proven (git, not PR descriptions)

One inaccuracy found & handled: PR #63 (SWE factory) was still open, not merged on main as the mission brief assumed. Merging devin/swe-factory-001 integrates it — its ancestry already contains current main. #63 can still merge cleanly to main whenever convenient (or be superseded by this PR).

Conflict resolution (1 conflict total)

package.json scripts — semantic union, not ours/theirs:

  • test: full vNext suite chain + swe-factory.test.mjs appended.
  • test:firestore: kept the vNext two-file form (firestore-emulator + firestore-vnext-emulator — superset of the factory side).
  • All six swe:* commands kept.
    AGENTS.md auto-merged (disjoint sections).

No R7 resurrection, no new features

Verification on final HEAD (9399726 + mission-artifact commit)

npm run verify:full — all green, one recorded run:

  • typecheck: 122 files OK
  • unit suites: all pass incl. vnext ×8, swe-factory 33 checks
  • build ✓
  • browser: 25 groups (vNext mission flow at 390px + 1280px incl. baseline→support→commit→retry→reload-resume)
  • Firestore emulator: both suites PASS (legacy progress + vNext evidence)
  • curriculum gate: 7 missions, 0 problems · pilot: replay determinism + learner isolation + no integrityFailures
  • Work Factory DONE gate passed on the merge HEAD

Stacked PRs #48/#51/#53/#56/#58/#60 remain open as review provenance until this lands.

Generated with Devin

Thunderkill016 and others added 30 commits September 28, 2026 17:08
…, planner (#42)

Greenfield learning core per docs/vnext/capability-model-v0.md — no UI,
no legacy coupling. The engine answers: what can the learner verifiably
do, on what evidence, did it survive delay, did it transfer, what next.

- capabilities.js: 12-capability graph across 5 modalities with an
  acyclic prerequisite check; Vietnamese risk probes attached per
  capability.
- evidence.js: append-only EvidenceEvent schema — 11 event types, typed
  context (mission/practiced/transfer/prompt family), observed vs
  self-report, per-support flags. answerBearing() is the single rule
  deciding what counts as help.
- projection.js: pure event-log → milestones → state derivation.
  Locks the doctrine: aided/self-reported success never earns
  INDEPENDENT; RETAINED needs a ≥24h unaided re-success; TRANSFERRED
  needs a novel promptFamily (replayed practiced prompts don't count);
  FLUENT = transfer across ≥2 families; modality-mismatched events earn
  nothing; replay is deterministic, ids dedupe on resync.
- risk-priors.js: 12 Vietnamese population priors — diagnostic probes
  only, learnerStateEffect is a constant, not a write path.
- planner.js: deterministic next action — resume → due retrieval →
  remediation (only for taught work; a baseline fail routes to teaching,
  not retry) → transfer → unaided attempt → probe/expose gated on
  INDEPENDENT prerequisites.
- tests/vnext.test.mjs: 9 groups pin every rule above plus the full
  vertical slice (baseline fail → input → retrieval → supported →
  feedback → retry → INDEPENDENT → RETAINED → TRANSFERRED → FLUENT).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
- only attempt-bearing event types may advance state; an exposure,
  support_use, or feedback record carrying outcome:success is contact,
  not performance
- canonical replay order is (occurredAt, id) — arrival order can no
  longer change the projection or the plan for equal timestamps
- projectLearnerState/planNext require learnerId; foreign-learner
  events are dropped before projection so a mixed log can never
  cross-contaminate
- a failed/partial due check routes to remediation instead of
  rescheduling delayed_retrieval forever
- FLUENT requires ≥2 novel transfer families AND ≥2 transfer successes
  measurably faster than the independent baseline
  (FLUENCY_LATENCY_RATIO); without latency data the ceiling is
  TRANSFERRED — two changed-context wins alone only prove breadth
- rehearsedPromptFamilies tracks every practiced-context family
  (aided or not); a support-rehearsed prompt can never be re-sold as
  a novel transfer context

Regression tests pin each invariant in tests/vnext.test.mjs §9.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Two evidence-honesty fixes from re-review:

- FLUENT is no longer auto-promotable. The latency-ratio rule was an
  uncalibrated invented threshold that rewarded answering faster —
  the doctrine requires hesitation + intelligibility + successful
  turns + repairs + stability. FLUENT stays in the enum/milestones as
  reserved but nothing sets it; TRANSFERRED is the v0 ceiling until a
  dedicated fluency-evidence contract exists.
- conditions.supportAllowed was dead data: an attempt could earn
  INDEPENDENT while using support the capability never permits.
  isIndependent now requires every used support kind to be within the
  capability's declared conditions. 'repeat_once' is only satisfied
  with provenance — support.repeatCount === 1; a bare repeat flag
  cannot distinguish once from many and fails the condition. Support
  fields gain repeatCount for this.

Regression tests pin both: repeat on a no-support capability caps at
SUPPORTED; a repeat_once capability accepts exactly one recorded
replay; even fast transfers stay TRANSFERRED.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
vnext: headless capability engine — graph, evidence, projection, planner (#42)
)

Teaching item ≠ transfer item ≠ assessment item — the engine now
enforces it structurally instead of trusting callers:

- contracts.js: MissionContract/TaskContract + validators. Missions
  check capability existence, prerequisite closure, task ownership,
  an eliciting path per target, and freshness collisions (a
  practiced family can never masquerade as fresh assessment/transfer).
  Tasks enforce purpose semantics: transfer needs declared changed
  dimensions + fresh_transfer family; assessment needs fresh_assessment
  family, zero allowed support, and no mid-attempt answer reveal.
- bind.js: bindAttempt/bindObservation are the only way learner-facing
  code mints EvidenceEvents. capabilityId/taskId/revision/modality/
  missionId/promptFamily/practicedOrTransfer/purpose/evaluation all
  derive from the contract — caller-supplied copies throw. eventType is
  constrained by purpose (an interaction task cannot emit checkpoint).
- effectiveAllowedSupport = capability ∩ task — a task may narrow,
  never broaden; violations are recorded honestly then demoted.
- projection: sticky support per attemptId (union across canonical
  order — a hint revealed mid-attempt permanently taints that attempt),
  evaluator-authority gate (only deterministic/human can award
  INDEPENDENT; self_report/asr/ai_llm cap at SUPPORTED), and the
  support-conditions check now uses the bound effective policy.
- fixtures.js: two complete headless missions — "meet a new person"
  (diagnostic → input → retrieval → supported → feedback → unaided →
  delayed → transfer → assessment) and "order a drink".
- validateMissionContent enforces declared language + injected budgets;
  no universal thresholds hard-coded.

tests/vnext-contracts.test.mjs pins all 15 required invariants.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
- evaluation.authority is contract-derived; a caller reporting a
  different authority is a binding error, never an upgrade path
- fresh_assessment binds its own 'assessment' context kind; only
  'transfer' context can feed TRANSFERRED, and assessed families are
  still rehearsed (cannot be re-sold as novel transfer)
- projection only credits contract-bound attempts: binding.purpose must
  exist and agree with eventType, so raw makeEvent() successes cap at
  SUPPORTED
- validateMission scopes evidence-path/freshness checks to taskIds and
  flags tasks that claim the mission without being declared
- validateMissionContent drops the capability-language whitelist:
  mission language must be declared in introduced/assumedKnown, and
  assumedKnown is structured {chunks,vocabulary,constructions}
- attempt-boundary support history keys on taskId::attemptId, so a
  reused attemptId cannot leak support across tasks

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ontracts

- projectLearnerState now takes the registered task list and re-derives
  every semantic field from it: taskId+revision, capabilityId, modality,
  purpose↔eventType, promptFamily, context kind, evaluation
  authority+contractId, and the recomputed effective support policy.
  A stamped binding is data, not proof — a raw makeEvent() carrying a
  forged binding can never verify and caps at SUPPORTED.
- Eliciting purposes require a real evaluation.contractId at
  construction (makeTask) AND at verification, so authority:
  deterministic + contractId: null can never mint independent evidence.
- planner passes the task registry through to the projection.
- tests: contract-test registry + core suite registers synthetic
  contracts mirroring binder output; new §13 pins forged-binding and
  missing-contractId cases.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
- The registry is keyed taskId@taskRevision so a v2 contract cannot
  overwrite v1 — historical evidence keeps verifying under the contract
  that produced it and replay stays stable. Duplicate id@revision
  registrations are rejected outright.
- verifyEventTask now requires the resolved task to pass validateTask()
  first — a hand-rolled invalid contract (e.g. purpose 'transfer' with
  zero changedDimensions) mints nothing, however closely an event's
  fields happen to match it.
- tests: revision coexistence, duplicate rejection, invalid-transfer
  regressions; the core suite's synthetic registry now meets the same
  purpose-specific requirements validateTask enforces.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
vnext: Mission/Task/Transfer/Assessment contracts + evidence binder (#45)
nextMissionTask() maps planner intents to concrete registered tasks
inside mission.taskIds — declared baseline diagnostics run first, then
expose/resume/retry/independent_attempt/delayed_retrieval/transfer map
to their exact purposes; a missing task yields blocked, never a silent
substitute. Assessment is gated on TRANSFERRED so it samples achieved
ability instead of standing in for it.

runMissionTrace() drives a scripted learner through selector → binder →
append-only evidence → projection, emitting an auditable per-step trace
(taskId, purpose, reason, before/after state).

Mission A gains the missing interact.ask_name phases (baseline
diagnostic, input, retrieval, remediation) so the full path
diagnostic fail → input → retrieval → supported → feedback → retry →
INDEPENDENT → 24h delayed → RETAINED → transfer → TRANSFERRED → fresh
assessment is selectable end-to-end and FLUENT stays unreachable.
Review round 2 orchestration fixes:

- Task identity is taskId@taskRevision end-to-end: selection returns
  taskRevision, mission ids resolve to the highest registered revision,
  trace lookup matches exact revision, duplicate id@revision throws.
- Consumption is verified-only: an event marks its task run only if it
  passes verifyEventTask against the registry — forged or mis-revisioned
  events can no longer hide an unrun baseline or input task.
- Mission integrity fails closed: a declared taskId absent from the
  registry or resolving to an invalid contract blocks selection instead
  of being dropped by filter(Boolean).
- expose/resume fallback is bounded to input/notice → retrieval /
  production / interaction — it can never drift into remediation,
  delayed_retrieval, transfer or assessment.
- EVENT_TYPES_FOR_PURPOSE gains input/notice → observation-only types so
  exposure tasks verify under the same purpose↔eventType contract.
…dence (#50)

Feasibility pilot for the evidence engine, not an efficacy study:
five scripted learners through the meet-a-new-person mission across
five deterministic sessions (day0, +25h, +50h, +75h, +100h).

- runPilotLearner: bounded runMissionTrace per session sharing one
  append-only log; harness stamps id/learnerId/occurredAt (learner
  isolation is the harness's job) and counts per-run service calls
  (ctx.call) so scripted outcomes are replay-deterministic.
- evaluateClaim: recomputes learnedByFlashday from PRIMITIVE verified
  evidence — milestones are outputs, never claim inputs. Baseline-pass
  → acquisitionSource PREEXISTING (never claimable); unresolved-
  contradiction rule replaces the brittle "no fail after independent";
  retained24h/retained72h horizons reported separately; milestone-vs-
  primitive mismatches surface as integrityFailures.
- mission-runner: a failed assessment is re-probed after remediation
  (latest checkpoint outcome ≠ success re-serves while TRANSFERRED
  holds) instead of dead-ending at first non-pass.
- tests/vnext-pilot: 10 checks — fast/dependent/forgetful/able/
  oscillator archetypes, claimRate over eligible claims, zero invalid
  promotions, forged-revision immunity, replay + isolation.

Research basis: issue #49 round 2 (feasibility criteria, unresolved-
contradiction claims, separate retention horizons, no vanity metrics).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Per R3 research (#49): engine defines what happened; policy defines
how much evidence is enough. Pedagogical thresholds are data.

- src/vnext/policy.js: LEARNING_POLICY_V1 frozen + registry,
  resolvePolicy fails closed on unknown versions/malformed objects,
  makePolicy for declared test variants. Thresholds: retention
  minLagMs, independent successfulUnaidedRetrievals/minDistinctSessions/
  minSpacingGapMs, remediation minConsecutiveFailures, claim require*
  flags + blockOnUnresolvedContradiction.
- projection: {policy} opt feeds retention lag; new engine fact
  consecutiveFailures (trailing fail/partial count); return carries
  policyVersion.
- planner: remediation routes on consecutiveFailures >=
  policy.remediation.minConsecutiveFailures (v1=1, same behavior).
- mission-runner/pilot-harness: policy plumbs through selection,
  trace and claim; evaluateClaim reads every threshold from policy
  and stamps policyVersion on the claim.
- tests/vnext-policy: registry fail-closed, frozen objects, same-
  events-different-lag retention, 2-strike remediation, version
  stamping, lenient-vs-strict claim divergence.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Per R3 research (#49): persistence lands before any UI — if a lost-ack
retry can double an attempt, the UI could mint INDEPENDENT by accident.

- src/vnext/persist.js: users/{uid}/vnext_events/{eventId} adapter on
  the cloud-compat fs surface. appendVnextEvents is per-event
  transactional — absent → create, identical content → dedupe,
  different content under the same id → hard conflict (evidence
  identity is never last-write-wins). loadVnextEvents returns the
  canonical (occurredAt, id) replay order. toEventDoc keeps both
  occurred_at (client-observed, drives lag) and recorded_at (server
  arrival) plus schema_version/policy_version/mission_run_id.
- firestore.rules: users/{uid}/vnext_events — create+read only; update
  AND delete denied even for the owner (immutable evidence, stricter
  than lesson_events); learner_id must equal the path uid — evidence
  for another learner is a forgery.
- tests/vnext-persist: fake-fs round-trip provenance, append → reload
  → replay parity (projection AND claim identical), exactly-once
  retry, id-conflict throw, out-of-order canonicalization, isolation.
- tests/firestore-vnext-emulator: real rules — owner CRUD, update/
  delete denied for owner, cross-owner + anon denied, learner_id
  forgery denied, malformed docs denied, e2e persist → reload →
  replay → identical claim.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
nextMissionTask → bindAttempt → append → projection, surfaced through
four screen types (intro / learn-input / reusable task / checkpoint).

- ui-session.js: headless session controller — selector-driven screens,
  deterministic attemptIds, support_use events + immutable snapshot on
  the attempt, prompt→committed→feedback phases, run persistence across
  reloads without a new missionRunId.
- evaluators.js: reusable deterministic contracts
  (eval.required_functions.v1, eval.choice.correct.v1) scoring real
  response shapes via mission-checks; fixtures now bind these contracts
  instead of hollow per-task ids.
- planner/selector: uncovered intents exclude per (capability, kind)
  via skipIntentFor — the cap stays on the surface for prerequisites
  and its other intents still route; only verified evidence consumes a
  task; novelty-free purposes may re-elicit, diagnostics stay
  single-sample.
- vnext/ page: Vietnamese-first copy describing observed evidence —
  no engine labels, no mastery claims; pre-commit support is recorded
  and stamps the attempt.
- tests: 13 ui-session checks + 2 browser groups covering baseline
  first, support before commit, stamped attempt, reload resume.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
)

R5 (issue #49) called for shipping the curriculum schema before scaling
content: a capability DAG, mission roles that separate claim-bearing
targets from carriers, and auditable prompt families.

- Capability ids namespaced by CEFR activity type:
  reception.listen.*, reception.read.*, production.speak.*,
  production.write.*, interaction.* — one migration, before any new
  missions depend on the names.
- Missions gain carrierCapabilities: rehearsed for retention/context
  but never owed transfer credit. Mission A now claims only the two
  R5-shaped targets (say_own_name + ask_name); say_own_name gets its
  missing delayed + held-out transfer tasks and checkpoint sample.
- contextSignature on every task: the family's auditable identity
  (cueTopology, setting, register, channel, interlocutorRole, ...).
  Canonical family ids pf.<cap>.<cue>.<setting>.<register>.<channel>.vN
  re-derive it, so an id that contradicts its signature fails.
- curriculum-checks.js is the authoring gate: DAG validity, per-target
  evidenceability (practiced + delayed + fresh_transfer + assessment
  sample, <=3 targets), family coherence (same id -> same signature),
  fake-novelty detection, and transfer deltas verified as real
  signature deltas against EVERY practiced family.
- Prerequisite closure now covers the whole declared surface (targets +
  carriers + support): an undeclared prereq deadlocks its dependents.
- Delayed re-checks deliberately reuse the rehearsed family — a delayed
  task on a novel family would measure transfer, not retention.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…roductions (#57)

Second authoring-gate pass from the GPT R6 design review:

- Canonical prompt-family ids gain a signature hash:
  pf.<cap>.<cue>.<setting>.<register>.<channel>.<sigHash8>.vN.
  The readable prefix stays human-inspectable; the fnv1a fingerprint
  covers the WHOLE contextSignature, so two families differing only in
  a non-id field (interlocutorRole, relationship, responseTopology,
  lexicalDomain) can no longer alias behind a .vN bump. Fixtures now
  derive ids via canonicalFamilyId(capId, sig) — id≡sig by construction.
- Missions are role-aware at introduction: targets owe a mandatory
  baseline probe; carriers rehearse opportunistically (input +
  eliciting practice, never a diagnostic, never fresh_transfer, never
  an assessment of their own); supports are demand-driven — undeclared
  rather than dead surface. Mission A drops the four carrier
  diagnostics and ask_repeat entirely; mission B's offer check becomes
  post-input retrieval practice.
- Surface cap: a mission declares at most 6 capabilities across all
  roles — beyond that the diagnostic/exposure phase explodes.
- interaction.order_drink broadens to interaction.request_item (same
  communicative function; drink vs food/goods is a context difference,
  not a capability difference).

Verify green: unit suites (curriculum, contracts, slice, ui-session,
pilot, policy, persist), browser 25 groups, Firestore emulator.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Five new A1 missions with full evidence packages (baseline diagnostic →
input → retrieval → interaction → delayed → transfer → checkpoint):
meet_at_a_time, complete_small_order, buy_small_item, find_a_place,
talk_about_self_family — 12 new capabilities in the CEFR namespaces.

Runner: uncovered intents on non-target caps (carriers, prerequisite
gates) are cross-mission backlog, not mission integrity failures —
stale prior-mission evidence no longer freezes later missions.
Curriculum gate: prerequisiteCapabilities are read-only DAG gates and
no longer count against the 6-capability active surface.

UI: all seven missions registered; mission page resolves the whole
curriculum task registry so carried prerequisite evidence verifies
(routing stays scoped to mission.taskIds).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…ms (#59)

Fake prerequisite edges removed: GPT R7's bar is that a hard edge is
only legitimate when no valid assessment exists that a learner could
pass downstream while failing upstream — sequencing belongs to pedagogy,
not the DAG. All capability prerequisites are now empty; curriculum
order moves to pedagogy.recommendedAfterMissions on the mission object.

- capabilities.js: all prereqs emptied; providesFunctions added on
  support caps (identify_spoken_number, understand_simple_choice,
  identify_basic_direction_term) as the substrate the future
  support-demand mechanism consumes.
- fixtures.js: mission declarations use targets/carriers only —
  paper-support declarations dropped per R7 (lazy support had no demand
  path); transfer signatures normalized to exactly the 3 designed
  dimensions (cueTopology, setting, interlocutorRole) so novelty claims
  carry ≤3 confounds; M6 gains a practiced landmark-retrieval task so
  follow_short_direction gets rehearsal evidence inside its own mission.
- tests: vertical-slice rewrite codifies the new semantic (baseline fail
  routes to teaching the capability itself, never a prereq walk); M3
  slice asserts target-first ordering, visible-but-non-fatal backlog,
  and lazy-support non-routing.
Repository-local orchestration for long autonomous missions: QUEUED →
RUNNING → VERIFYING → DONE/FAILED/BLOCKED, one non-terminal mission at
a time, state in plain artifacts under missions/<id>/.

- scripts/swe.mjs: status/start/checkpoint/verify/finish/resume.
  swe:start refuses on dirty tree or a second active mission.
  Verification commands are allowlisted shapes (npm run <script>,
  node tests|scripts/<file>) — a mission file can never smuggle shell.
  swe:finish DONE requires the last verify green on current HEAD.
  Checkpoint ingest stamps git state and enforces all 10 sections.
  swe:resume rebuilds context from mission + latest checkpoint + git.
- missions/TEMPLATE.md + README.md: authoring contract + lifecycle doc.
- tests/swe-factory.test.mjs: 21 checks in a throwaway git repo —
  dirty-tree refusal, one-active-mission, checkpoint validation,
  verify gating incl. HEAD-moved re-verify, red-verify cannot DONE,
  blocked→resume, allowlist rejection, terminal immutability.
- npm scripts swe:* wired; test wired into npm test.

Fixed en route: git status --porcelain output must not be trimmed —
the first line's leading space is the unstaged-worktree status column;
trimming corrupted path extraction ('tests/x' → 'ests/x').
001-factory-smoke: start → checkpoint → verify (PASS) → finish DONE,
REPORT.md generated. Zero product changes; files-changed-vs-start = 0.

002-resume-sim: start → checkpoint --blocked → swe:resume printed the
full context packet and unblocked → verify (PASS) → DONE. Proves a
fresh session reconstructs state from mission + checkpoint + git alone.
…ontract (#62)

Two false-success paths closed (external review findings):

- Finding 1 (confirmed): DONE could land while tracked product changes
  sat uncommitted — verify ran on the dirty tree and finish only compared
  SHAs, so the report's startSha..HEAD diff hid the work. swe:finish
  DONE now requires a clean tree per the same policy as swe:start;
  --result failed stays unconditional (abandon never traps work).

- Finding 2 (confirmed): mission.md was re-read per command, so a
  running mission could rewrite objective/criteria/verification —
  committed or not — and still finish DONE. swe:start now pins the
  mission.md sha256 into state.json; verify/finish/resume re-hash and
  fail closed on drift. Restoring original bytes restores operation.
  There is deliberately no V1 amendment workflow.

- README: allowlist wording corrected — it constrains mission-specified
  verification commands only, not the agent's own tool access
  (reliability guard, not a sandbox).

Tests: 10 new checks — dirty-tree refusal (tracked + unsafe untracked),
commit-then-reverify-then-done, uncommitted/committed mission.md drift
detection at all three guarded commands, restore-definitions-recovers,
hash persisted at start.
@Thunderkill016
Thunderkill016 merged commit 27574cb into main Sep 30, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant