Skip to content

VNEXT-CURRIC-001: Capability curriculum DAG beyond the two fixture missions #57

Description

@Thunderkill016

Research-driven curriculum expansion (GPT R3 on #49, workstream B after
persistence + minimal UI).

Problem

The engine + mission runner + honest UI are proven end-to-end on the
meet-a-person slice. The capability graph currently has ~12
capabilities and 2 fixture missions — hand-placed, not a designed
curriculum. To teach toward A1 coverage we need:

  • A capability DAG that maps communicative A1 outcomes (CEFR/GSE-anchored)
    to observable elicitations, not a flat list.
  • Mission families that cover the DAG with the task phases the contracts
    require (diagnostic → input → retrieval → interaction → remediation →
    delayed → transfer → assessment).
  • Transfer contexts that are meaningfully different per capability —
    held-out prompt families encoded at authoring time, never improvised.
  • Vietnamese learner risk priors fed by real pilot data, not guesses.

Constraints (do not violate)

  • Exposure is not evidence; supported ≠ unaided; completion ≠ mastery.
  • Every eliciting task needs a real deterministic evaluator contract.
  • Mission content must satisfy validateTask + content budgets.
  • No new planner semantics just to fit content — if a capability can't
    be elicited honestly, it doesn't ship.

Design input

GPT R3 on #49: near-transfer for ask_name = interlocutor +
conversational lead-in change; retention and transfer probes must stay
separate; INDEPENDENT needs spaced unaided successes (policy-level).
Next GPT round (R5) will produce the curriculum skeleton to encode.

Out of scope

Auth changes, AI tutor, streaks, production deploy, UI polish beyond
the four existing screen types.

Activity

  1. Thunderkill016 commented on Sep 30, 2026

    @Thunderkill016
    OwnerAuthor

    Progress — PR #58: curriculum schema + gate landed on devin/vnext-curric-001 (stacked on #56).

    Shipped

    • Namespace migration: reception.listen.*, reception.read.*, production.speak.*, production.write.*, interaction.* — done once now so new missions/families never re-migrate.
    • Mission roles: targetCapabilities (claim-bearing, ≤3, prefer 2) vs carrierCapabilities (retention/context, no transfer owed) vs supportCapabilities (auxiliary). Mission A targets = say_own_name + ask_name; carriers get baseline/input/retrieval evidence only.
    • Canonical prompt families: pf.<cap>.<cueTopology>.<setting>.<register>.<channel>.vN + contextSignature on every task — novelty is now auditable data, not prompt-text vibes.
    • curriculum-checks.js authoring gate: DAG validity + per-target evidenceability (practiced/delayed/transfer/assessment-sample) + family coherence + fake-novelty detection + transfer deltas verified against every practiced signature.
    • say_own_name full claim path added: delayed.name, transfer.name (educational/name_check context), checkpoint capabilitySample + state_own_name in required functions.
    • Honest semantics: delayed re-checks reuse rehearsed families (novel-family delayed would measure transfer, not retention); validateMission prereq closure now covers the whole surface.
    • Tests: new vnext-curriculum.test.mjs in npm test; slice walks 17 steps with both targets; negative coverage for every gate rule.
    • All gates green: npm test, typecheck (120 files), build, test:browser (25 groups), test:firestore.

    PR: #58

    Next (after merge)

    Seed missions 3–7 per R5 skeleton (time/food/price/directions/self-description) + reception.* foundation capabilities — each new target ships with the full evidence package or the gate fails it.

  2. Thunderkill016 commented on Sep 30, 2026

    @Thunderkill016
    OwnerAuthor

    Wave 2 pushed — R6 semantics landed (27f98da)

    Second authoring-gate pass on devin/vnext-curric-001, implementing the
    GPT R6 adjustments (full research log: #49):

    Schema

    • Injective family ids — pf.<cap>.<cue>.<set>.<reg>.<chan>.<sigHash8>.vN.
      R6 found the non-injective hazard: two signatures differing only in
      interlocutorRole/relationship/responseTopology/lexicalDomain
      could previously share a readable prefix and differ only in vN,
      conflating authoring revision with family distinction. sigHash8 =
      fnv1a of the canonicalized FULL signature; signatureHash() +
      canonicalFamilyId() exported from contracts.js; fixtures build ids
      through it so id≡sig by construction. Gate: hash segment verified
      against the task's actual signature (mismatch → reject).

    Role-aware semantics

    • Planner takes roles (targets/carriers/supports) + rolesSource —
      introduction routing only, never evidence semantics.
      • target → mandatory baseline diagnostic probe.
      • carrier → no baseline; rehearses opportunistically (exposure +
        practice only). Gate now rejects diagnostics, assessments, and
        fresh_transfer tasks on carriers.
      • support → demand-driven; carriers vs supports distinguished so
        undeclared supports don't create dead surface.
    • Gate adds: diagnosticCoverage (every target owes a diagnostic),
      forbiddenClaimTasks (carriers), surfaceBudget (≤6 declared caps).

    Mission fixes

    • Mission A slimmed: dropped 4 carrier diagnostics + ask_repeat
      (no demand mechanism — GPT's call). Slice goes 15→16/17 steps.
    • Mission B: assessment_offer demoted to post-input retrieval.offer
      (it wasn't gated on transfer → it was off-invariant rehearsal,
      now correctly typed; baselinePassed no longer pretends offer
      covered the carrier).
    • interaction.order_drink → interaction.request_item
      (request_item matcher broadened for generic service items).

    Tests updated: slice (17-step happy path, negative-path asserts),
    ui-session (carrier tasks in path, progress snapshot counts 8 surface
    caps), curriculum gate (+2 checks covered), contracts (hash-aware
    parse expectations), persist (conflict now mutates a real attempt event —
    events[4] became an exposure event), pilot/policy (input tasks observe
    instead of attempt), browser vnext drive.

    Verify: npm run verify (typecheck+tests+build) ✓, test:browser
    25 groups ✓, test:firestore PASS ✓.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions