Repository navigation
vnext: headless capability engine — graph, evidence, projection, planner (#42) - #43
Conversation
…, planner (#42) Greenfield learning core per docs/vnext/capability-model-v0.md — no UI, no legacy coupling. The engine answers: what can the learner verifiably do, on what evidence, did it survive delay, did it transfer, what next. - capabilities.js: 12-capability graph across 5 modalities with an acyclic prerequisite check; Vietnamese risk probes attached per capability. - evidence.js: append-only EvidenceEvent schema — 11 event types, typed context (mission/practiced/transfer/prompt family), observed vs self-report, per-support flags. answerBearing() is the single rule deciding what counts as help. - projection.js: pure event-log → milestones → state derivation. Locks the doctrine: aided/self-reported success never earns INDEPENDENT; RETAINED needs a ≥24h unaided re-success; TRANSFERRED needs a novel promptFamily (replayed practiced prompts don't count); FLUENT = transfer across ≥2 families; modality-mismatched events earn nothing; replay is deterministic, ids dedupe on resync. - risk-priors.js: 12 Vietnamese population priors — diagnostic probes only, learnerStateEffect is a constant, not a write path. - planner.js: deterministic next action — resume → due retrieval → remediation (only for taught work; a baseline fail routes to teaching, not retry) → transfer → unaided attempt → probe/expose gated on INDEPENDENT prerequisites. - tests/vnext.test.mjs: 9 groups pin every rule above plus the full vertical slice (baseline fail → input → retrieval → supported → feedback → retry → INDEPENDENT → RETAINED → TRANSFERRED → FLUENT). Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Thunderkill016
left a comment
There was a problem hiding this comment.
ARCHITECT REVIEW — NOT ACCEPTED YET
The greenfield separation and overall direction are correct, and CI is green. However, this cannot merge yet because the current projection/planner can over-claim capability state or get stuck in a bad learning loop.
Blockers
-
Only real attempt events may advance capability state.
projection.jsdefinesATTEMPT_TYPES, butisSuccess(e)ignores it. Today anexposure,feedback, orsupport_useevent withattempt.outcome:'success'can become SUPPORTED/INDEPENDENT/RETAINED/TRANSFERRED.- Gate state-advancing success on
ATTEMPT_TYPES.has(e.eventType). - Add regression tests for success-looking non-attempt events.
- Gate state-advancing success on
-
Replay is not actually arrival-order deterministic for equal timestamps.
Events are sorted byoccurredAtthen original array index. If two same-timestamp events arrive in a different order after sync,lastAttemptOutcome(and therefore planner behavior) can differ.- Use a stable domain key (e.g.
occurredAt, then event id) rather than input index. - Add a test that reversed arrival order produces identical projection and plan.
- Use a stable domain key (e.g.
-
Learner isolation is missing from the projection contract.
EvidenceEventcontainslearnerId, butprojectLearnerState(events,...)currently merges all learners in the provided log.- Require a learner id and filter/validate events, or explicitly reject a mixed-learner log.
- Add a cross-learner isolation regression test.
-
Failed delayed retrieval can loop forever.
Planner priority checks due retrieval before remediation. After an overdue delayed retrieval fails,lastIndependentSuccessAtremains old, so the capability is still due and the planner selects anotherdelayed_retrievalinstead of feedback/self-repair.- A failed/partial due check must route to remediation/retry before rescheduling another delayed check.
- Add a regression test.
-
FLUENT = two transfer prompt familiesoverclaims fluency.
Our vNext doctrine defines fluency as repeated transfer-capable performance with lower hesitation/stable intelligibility while preserving meaning. Two changed-context successes prove broader transfer, not fluency.- Do not auto-promote to FLUENT from transfer-count alone.
- For this PR, FLUENT may remain unreachable until a dedicated fluency evidence contract exists, or add an explicit measurable fluency rule.
- Update the vertical-slice test accordingly.
-
Transfer novelty must be measured against practiced exposure, not only previously independent prompt families.
provenPromptFamiliesrecords a practiced family only after independent success. If a learner previously saw/practiced a prompt family with support, a latertransfer_attemptusing that same family can incorrectly count as novel transfer.- Track practiced/encountered prompt families separately from proven families.
- Same practiced family must never become transfer merely because earlier practice was aided.
Non-blocking cleanup
CAPABILITY_STATESimport in planner is unused.- Re-check risk-probe mappings against the actual language feature of each capability before expanding them.
Do not start UI work. Fix these invariants, add the regression tests, get CI green on the new head, then re-review.
- only attempt-bearing event types may advance state; an exposure, support_use, or feedback record carrying outcome:success is contact, not performance - canonical replay order is (occurredAt, id) — arrival order can no longer change the projection or the plan for equal timestamps - projectLearnerState/planNext require learnerId; foreign-learner events are dropped before projection so a mixed log can never cross-contaminate - a failed/partial due check routes to remediation instead of rescheduling delayed_retrieval forever - FLUENT requires ≥2 novel transfer families AND ≥2 transfer successes measurably faster than the independent baseline (FLUENCY_LATENCY_RATIO); without latency data the ceiling is TRANSFERRED — two changed-context wins alone only prove breadth - rehearsedPromptFamilies tracks every practiced-context family (aided or not); a support-rehearsed prompt can never be re-sold as a novel transfer context Regression tests pin each invariant in tests/vnext.test.mjs §9. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
Round 2 — all 6 blockers addressed in 1. Attempt-only state advancement — 2. Order-free replay — canonical sort is 3. Learner isolation — 4. No due-loop — planner rule 2 skips a capability whose latest attempt is 5. Honest fluency — 6. Practiced ≠ novel — Cleanup: unused |
Thunderkill016
left a comment
There was a problem hiding this comment.
ARCHITECT RE-REVIEW — STILL NOT READY
Head reviewed: 2df6c5bd53a8066ee9b9f638ef6dd691230888d5
CI: verify:full green.
The fixes for blockers 1, 2, 3, 4 and 6 are accepted:
- state advancement is attempt-only;
- replay order is canonical;
- learner isolation is explicit;
- failed delayed retrieval routes to remediation;
- rehearsed prompt families can no longer be re-sold as novel transfer.
Two evidence-honesty blockers remain.
A. FLUENT is still over-claimed
The new rule is measurable, but the threshold is invented:
FLUENCY_LATENCY_RATIO = 0.8
Two transfer successes faster than 80% of the first-independent latency are automatically promoted to FLUENT.
This does not satisfy our own vNext doctrine:
- capability spec: FLUENT requires lower hesitation and stable intelligibility while preserving meaning;
- Vietnamese-first doctrine: fluency should consider hesitation, response latency, successful turns, repairs, intelligibility and stability across repetitions;
- explicit rule: do not reward simply speaking faster.
The 0.8 ratio has no calibration/evidence behind it, and the same mechanism applies to every modality.
For #43, the safe fix is simpler:
- keep
FLUENTin the enum/schema if desired; - do not auto-promote to FLUENT yet;
- ceiling at
TRANSFERRED; - create a later dedicated fluency contract when we have measurable multi-signal evidence and calibration.
Issue #42's first implementation target ends at changed-context transfer anyway, so FLUENT is not required for this PR.
B. Capability conditions are currently dead data for evidence validity
Capability.conditions.supportAllowed exists in the contract, but projection never checks it.
All current capabilities inherit:
supportAllowed: []
Yet the test explicitly asserts:
repeat: true -> INDEPENDENT
because repeat is globally treated as non-answer-bearing.
That means an attempt can earn INDEPENDENT even when the capability's own conditions say no support is allowed. This breaks the core definition of “valid task under defined conditions.”
Required fix:
- make independent evidence respect capability conditions;
- at minimum, support used must be allowed by that capability's conditions;
- if we want
repeat_once, evidence needs enough provenance to distinguish one repeat from unlimited repeats, otherwise do not claim that condition.
Add regression tests:
- repeat used on a capability with no repeat support allowed cannot earn INDEPENDENT;
- allowed repeat condition can earn valid independent evidence only within its declared limit/contract.
Non-blocking
The old risk-probe cleanup is still only partially addressed: checking appliesTo ∩ modality is not the same as checking linguistic relevance. Example: vn.theta_eth is attached to listen.identity_question_basic, but the listed target chunks do not obviously contain /θ/ or /ð/. Keep this for the content-validation workstream; it does not need to block #43.
Once A+B are fixed and exact-head CI is green, I will re-review and merge.
Two evidence-honesty fixes from re-review: - FLUENT is no longer auto-promotable. The latency-ratio rule was an uncalibrated invented threshold that rewarded answering faster — the doctrine requires hesitation + intelligibility + successful turns + repairs + stability. FLUENT stays in the enum/milestones as reserved but nothing sets it; TRANSFERRED is the v0 ceiling until a dedicated fluency-evidence contract exists. - conditions.supportAllowed was dead data: an attempt could earn INDEPENDENT while using support the capability never permits. isIndependent now requires every used support kind to be within the capability's declared conditions. 'repeat_once' is only satisfied with provenance — support.repeatCount === 1; a bare repeat flag cannot distinguish once from many and fails the condition. Support fields gain repeatCount for this. Regression tests pin both: repeat on a no-support capability caps at SUPPORTED; a repeat_once capability accepts exactly one recorded replay; even fast transfers stay TRANSFERRED. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
|
Round 3 — both evidence-honesty blockers fixed in A. FLUENT auto-promotion removed. B. Support conditions enforced.
New §10 regression block:
Note: no shipped capability claims |
Thunderkill016
left a comment
There was a problem hiding this comment.
ARCHITECT ACCEPTED — FINAL
Reviewed exact head 1bcf36f92287f1b31b63b6a2067c6f024e4abca0.
The two final evidence-honesty blockers are resolved:
- v0 has no automatic path to FLUENT; TRANSFERRED is the ceiling until a calibrated multi-signal fluency contract exists.
- independent evidence now respects capability support conditions, with explicit provenance for repeat_once.
Previously accepted invariants remain intact: attempt-only credit, canonical replay, learner isolation, delayed-failure remediation, modality isolation, practiced-vs-transfer separation, and diagnostic-only VN priors.
Exact-head CI is green. Accepted for merge into rebuild/vnext.
Summary
Implements issue #42 headless, per
docs/vnext/capability-model-v0.md. No UI, no auth, no legacy coupling —src/vnext/is greenfield. The repo can now answer in code: what can this learner verifiably do, on what evidence, did it survive delay, did it transfer, and what happens next.capabilities.js— 12 capabilities × 5 modalities, acyclic prerequisite graph, VN risk-probe referencesevidence.js— append-onlyEvidenceEvent(11 types);answerBearing()is the one rule deciding "helped"projection.js— pure replay → per-capability milestones + state (NOT_SEEN…FLUENT)risk-priors.js— 12 VN population priors, diagnostic-only (learnerStateEffect: 'none_without_observed_evidence')planner.js— deterministic priority: resume → due retrieval → remediation → transfer → independent attempt → probe/expose (prereq-gated)Locked invariants (all pinned in
tests/vnext.test.mjs)SUPPORTED, neverINDEPENDENT(repeatis not answer-bearing)RETAINEDrequires an unaided success ≥RETENTION_DELAY_MS(24h) after first independenceTRANSFERREDrequirespracticedOrTransfer: 'transfer'AND a novelpromptFamily— replaying the practiced prompt doesn't countFLUENT= transfer across ≥2 distinct prompt familiesmodality≠ capability modality earn zero credit — speaking can never mark listening learneddiagnostic_probeactions; a fresh learner's state is untouchedfailroutes to teaching (prereq), notretryTest plan
node tests/vnext.test.mjs— 9 groups incl. the full vertical slice (baseline fail → input → retrieval → supported → feedback → retry → INDEPENDENT → RETAINED → TRANSFERRED → FLUENT)npm run verify— typecheck 98 files + all node suites + build ✅Generated with Devin