Parent: #39
Depends on: #42, #44, #45 (completed)
Blocks: Vietnamese beginner pilot and learner-facing UI
Goal
Prove one complete FlashDay vNext learning loop headlessly using the real Mission/Task/Evidence/Projection/Planner contracts.
Primary slice:
Mission: meet a new person
Target capability: interact.ask_name
The slice must demonstrate:
baseline diagnostic fail
→ comprehensible input
→ retrieval
→ supported interaction
→ feedback
→ self-repair / retry
→ clean independent attempt
→ delayed retrieval (24h+)
→ changed-context transfer
→ fresh assessment
No UI in this issue.
Why this exists
Unit tests currently prove individual invariants. They do not yet prove that the complete system can drive one learner through a coherent learning journey without a human manually choosing every next task.
This issue adds the minimum orchestration needed to answer:
Given the learner's evidence log, this mission's task contracts, and the current time, which exact task should run next, why, and what state transition is expected?
1. Mission runner / task selector
Implement a pure deterministic layer under src/vnext/.
Suggested API:
nextMissionTask({
learnerId,
mission,
tasks,
capabilities,
events,
riskPriors,
now
}) -> {
status,
taskId,
capabilityId,
purpose,
reason
}
The selector must:
- consume projected learner state, planner decision, MissionContract and TaskContract;
- return a concrete registered Task, not just a capability;
- never invent a task/prompt at runtime;
- be deterministic for the same inputs;
- explain why the task was selected.
2. Planner → task mapping
Map planner intents to compatible task purposes, conservatively:
- diagnostic_probe → diagnostic
- expose/resume → input/notice or pending mission phase
- retry → remediation
- independent_attempt → retrieval/production/interaction, unaided-compatible
- delayed_retrieval → delayed_retrieval
- transfer → transfer
- checkpoint → assessment
Do not map a planner intent to a task whose purpose cannot support that intent.
If there is no valid task, return an explicit blocked state with a reason. Never silently substitute a different semantic purpose.
3. Slice-specific task coverage
Mission A must contain a coherent path for interact.ask_name itself.
Add/adjust authored tasks if necessary so the target capability has:
- baseline diagnostic task;
- input/context task;
- retrieval task;
- supported interaction/remediation path;
- clean independent interaction task;
- delayed retrieval task;
- fresh transfer task;
- fresh assessment task.
Every eliciting task keeps a named evaluator contract.
4. Attempt lifecycle
Use bindAttempt / bindObservation only.
The slice test must append real EvidenceEvents in sequence. Do not directly mutate learner state.
Support provenance must survive retry boundaries according to #45.
A clean independent attempt must use a fresh attempt boundary when the contract permits it; prior help is not erased.
5. Time semantics
Use explicit deterministic timestamps.
At minimum:
- immediate same-session success does not count as retained;
- delayed retrieval occurs only after
RETENTION_DELAY_MS;
- transfer happens only after retention in this slice;
- assessment is separate from transfer.
Do not use wall-clock time in tests.
6. Expected state checkpoints
The slice must assert learner state after meaningful stages:
baseline fail -> EXPOSED
supported success -> SUPPORTED
clean unaided success -> INDEPENDENT
24h+ delayed success -> RETAINED
fresh transfer success -> TRANSFERRED
fresh assessment -> remains TRANSFERRED
FLUENT remains unreachable.
7. Negative paths
Regression tests must include:
- baseline pass skips unnecessary teaching for that capability;
- support-aided retry does not become independent;
- failed delayed retrieval routes to remediation;
- missing transfer task returns blocked instead of reusing practiced task;
- fresh assessment success cannot substitute for transfer;
- task revision mismatch cannot be selected/credited;
- planner/selector replay is deterministic;
- foreign learner events do not affect the selected task.
8. Trace output
Provide a headless trace helper for tests/dev diagnostics:
[
{ step: 1, taskId, purpose, reason, beforeState, afterState },
...
]
This is diagnostic output, not learner UI.
The trace must make it possible to audit why the learner moved from one state to another.
9. Definition of done
A single automated test drives a synthetic Vietnamese beginner through the entire interact.ask_name slice using:
- MissionContract;
- TaskContract;
- registered revision-aware task catalog;
- binder;
- append-only evidence;
- projection;
- planner;
- deterministic concrete-task selector.
No direct state mutation. No fake binding. No UI.
The final trace must end at:
interact.ask_name = TRANSFERRED
with a separate fresh assessment recorded, while:
Out of scope
- browser UI;
- Firebase/auth;
- speech recognition implementation;
- LLM tutor;
- production content library;
- streak/XP;
- persistence;
- pilot recruitment.
Once this is merged, the next workstream is a small Vietnamese beginner pilot harness, then minimal UI.
Parent: #39
Depends on: #42, #44, #45 (completed)
Blocks: Vietnamese beginner pilot and learner-facing UI
Goal
Prove one complete FlashDay vNext learning loop headlessly using the real Mission/Task/Evidence/Projection/Planner contracts.
Primary slice:
The slice must demonstrate:
No UI in this issue.
Why this exists
Unit tests currently prove individual invariants. They do not yet prove that the complete system can drive one learner through a coherent learning journey without a human manually choosing every next task.
This issue adds the minimum orchestration needed to answer:
1. Mission runner / task selector
Implement a pure deterministic layer under
src/vnext/.Suggested API:
The selector must:
2. Planner → task mapping
Map planner intents to compatible task purposes, conservatively:
Do not map a planner intent to a task whose purpose cannot support that intent.
If there is no valid task, return an explicit blocked state with a reason. Never silently substitute a different semantic purpose.
3. Slice-specific task coverage
Mission A must contain a coherent path for
interact.ask_nameitself.Add/adjust authored tasks if necessary so the target capability has:
Every eliciting task keeps a named evaluator contract.
4. Attempt lifecycle
Use
bindAttempt/bindObservationonly.The slice test must append real EvidenceEvents in sequence. Do not directly mutate learner state.
Support provenance must survive retry boundaries according to #45.
A clean independent attempt must use a fresh attempt boundary when the contract permits it; prior help is not erased.
5. Time semantics
Use explicit deterministic timestamps.
At minimum:
RETENTION_DELAY_MS;Do not use wall-clock time in tests.
6. Expected state checkpoints
The slice must assert learner state after meaningful stages:
FLUENT remains unreachable.
7. Negative paths
Regression tests must include:
8. Trace output
Provide a headless trace helper for tests/dev diagnostics:
This is diagnostic output, not learner UI.
The trace must make it possible to audit why the learner moved from one state to another.
9. Definition of done
A single automated test drives a synthetic Vietnamese beginner through the entire
interact.ask_nameslice using:No direct state mutation. No fake binding. No UI.
The final trace must end at:
with a separate fresh assessment recorded, while:
Out of scope
Once this is merged, the next workstream is a small Vietnamese beginner pilot harness, then minimal UI.