Skip to content

VNEXT-RESEARCH-003: A1 outcome calibration, pilot design and curriculum research #49

Description

@Thunderkill016

Parent: #39
Related: #44 (research round 2, closed), #45 (contract layer, merged), #47 / PR #48 (headless vertical slice)

Goal

Standing research collaboration on the learning-methodology and content side of FlashDay vNext. The headless engine now exists (capability/evidence/contract layers are merged), so the open questions are no longer architectural — they are empirical and pedagogical:

What must a synthetic or real Vietnamese beginner demonstrably DO for us to claim A1, and what content/method gets them there?

This issue is the discussion channel. Post findings as comments; Devin consumes them into contract/code changes and asks follow-ups here. The loop runs until the web product is complete — this is a program, not a one-shot report.

Architecture already fixed (do not re-litigate)

  • Capability state: NOT_SEEN → EXPOSED → SUPPORTED → INDEPENDENT → RETAINED → TRANSFERRED → FLUENT (FLUENT unreachable in v0 — needs a calibrated contract).
  • Teaching ≠ Transfer ≠ Assessment. Fresh prompt families for transfer and assessment.
  • Support provenance is sticky per taskId@taskRevision attempt; answer-reveal can never launder into independent evidence.
  • Only deterministic/human evaluators mint INDEPENDENT; asr/ai_llm/self_report cap at SUPPORTED.
  • Selector is deterministic, explainable, revision-aware, fail-closed on mission integrity.
  • No cloud/API-key AI is allowed in the product direction; on-device only.

Research questions (in priority order)

R1 — A1 outcome calibration

  • Map CEFR A1 descriptors + GSE A1 learning objectives onto observable elicitation types. Which can-dos are actually assessable at A1 for: meeting people, ordering, directions, personal info, daily routine?
  • What does the evidence base say a credible 'A1 achieved' claim requires (breadth of functions × independence × retention × transfer)? Where do typical apps overclaim?
  • What would a defensible FLUENT contract look like at A1 scale — hesitation, intelligibility, successful turns, repairs, stability — with measurable thresholds?

R2 — Vietnamese beginner pilot design

  • For a small headless pilot (no UI): what does a session look like, what cadence, how many capabilities in flight?
  • What is the minimum evidence package per capability before we claim the loop worked?
  • Failure taxonomy: which failure patterns should route to which remediation? Feed src/vnext/risk-priors.js — Vietnamese-specific phonological/grammatical error priors with sources.

R3 — Content + curriculum design inside the contract

  • How to author context-rich tasks that satisfy TaskContract constraints (promptFamily, freshness, transfer.changedDimensions)? Concretely: what are the meaningful transfer dimensions for 'meet a new person', 'order a drink'?
  • Multi-session curriculum ordering: which missions/capabilities unlock which — is the current prerequisite graph defensible pedagogy or does the research suggest a different DAG?
  • Retrieval dosage: how many exposures / retrievals / varied contexts does the literature support per phase at A1? What are sane thresholds for the planner (support decay, remediation triggers, delayed windows)?

R4 — Product/UI research (downstream of pilot)

  • Minimal UI patterns for input → retrieval → feedback → retry that preserve evidence honesty (no fake 'I got it' buttons, explicit support provenance).
  • What do the best beginner products show vs. hide during the learning loop? Reference systems only — product claims are not learning evidence.

Method requirements

Same hierarchy as #44: meta-analysis → controlled studies → transparent efficacy data → official docs → product claims (labeled as such). Vietnamese-learner evidence explicitly sourced where used. Every recommendation must name which engine module or content contract it affects — findings that cannot touch the system are out of scope.

Definition of done (per round)

A round closes when its findings are concrete enough that Devin can either (a) encode them into a contract, planner threshold, risk prior or content fixture, or (b) record a reasoned rejection. Open rounds stay open; the issue stays open until the web product ships.

No learner-facing UI is built inside this issue.

Activity

  1. Thunderkill016 commented on Sep 29, 2026

    @Thunderkill016
    OwnerAuthor

    Kickoff from Devin — trạng thái hiện tại để bắt đầu trao đổi:

    Engine đã chứng minh headless end-to-end (PR #48 đang review): một synthetic learner đi hết baseline fail → input → retrieval → supported → feedback → retry → INDEPENDENT → +24h delayed → RETAINED → transfer context mới → TRANSFERRED → fresh assessment riêng, FLUENT unreachable.

    Điểm nghẽn tiếp theo không phải kiến trúc mà là calibration: các ngưỡng planner hiện là constants tự đặt (RETENTION_DELAY_MS=24h, support policy rỗng cho mọi capability, remediation trigger = lastOutcome fail/partial). Cần evidence để gán con số có cơ sở.

    Câu hỏi mở đầu (R2 + R3):

    1. Pilot design — nếu chạy headless pilot với ~5 người Việt mới học trên mission 'meet a new person': cấu trúc buổi học, số capability song song, và gói evidence tối thiểu/capability trước khi claim loop đúng là gì? Cần nghiên cứu nào backing câu trả lời?

    2. Transfer dimensions — contract yêu cầu transfer task đổi ≥1 dimension (partner/topic/channel/setting...). Cho 'interact.ask_name', những dimension nào là meaningful context change theo nghiên cứu transfer-of-learning (không chỉ đổi surface word)? Ordering nào nên áp cho multi-mission curriculum?

    3. Dosage — literature về retrieval practice/spacing ở trình độ absolute-beginner: bao nhiêu lần retrieval thành công / bao nhiêu context variation trước khi delayed test có nghĩa? Số này sẽ thành planner thresholds thật thay constants tự đặt.

    Trả lời bằng comment trên issue này — mỗi finding cố gắng kèm: nguồn (level trong hierarchy), module bị ảnh hưởng (planner/contracts/fixtures/risk-priors), và 'adopt/adapt/reject/test' suggestion. Tao sẽ consume vào code/contract và hỏi tiếp tại đây.

  2. Thunderkill016 commented on Sep 29, 2026

    @Thunderkill016
    OwnerAuthor

    Research round 1 — GPT response (via Playwright exchange, model: ChatGPT 'Cao' thinking)

    Source: live ChatGPT session, prompt = kickoff comment above (3 questions: pilot design, transfer dimensions, dosage). Response captured verbatim below — evaluate each finding against the contract layer before adopting.


    Tôi đã ưu tiên meta-analysis → controlled studies → official framework; không dùng product docs vì 3 câu này đã có bằng chứng học thuật tốt hơn. Điểm cần giữ kỷ luật nhất: nhiều ngưỡng dưới đây là FlashDay decision rule để test, không phải “con số khoa học đã chứng minh”.

    Kết luận ngắn

    Với FlashDay hiện tại, tôi sẽ chốt pilot theo hướng này:

    n≈5 = feasibility/evidence-engine pilot, không phải efficacy study. Đừng dùng nó để kết luận FlashDay “dạy hiệu quả”; dùng để xem state machine có phân biệt đúng supported / independent / retained / transferred hay không. Pilot nhỏ vốn không phù hợp để hypothesis-test hiệu quả.
    DOI
    +1
    Trong một mission, chỉ để 2 target capabilities chạy song song, tối đa 3 nếu capability thứ ba chỉ là prerequisite/observational. Đây là quyết định instrumentation của FlashDay, không phải ngưỡng từ meta-analysis.
    Với mỗi capability, trước khi cho INDEPENDENT, yêu cầu 2 successful retrievals không scaffold, có một khoảng chen giữa; đừng yêu cầu 5–7 chỉ vì muốn “chắc”. Sau đó test lại ở +24h không feedback trước response.
    RETAINED nên lưu rõ retention_horizon = 24h. Một pass sau 24h không chứng minh long-term retention.
    Transfer không phải “đổi John → Mary”. Với ask_name, thay interlocutor + conversational lead-in là transfer gần tốt nhất để pilot. Thay channel text→voice là meaningful nhưng đã thay luôn skill/modality, nên để sau.
    Retention test và transfer test phải tách nhau. Nếu +24h vừa đổi context vừa test, khi learner fail bạn không biết họ quên hay không generalize được.

    1. PILOT DESIGN — 5 người, mission “meet a new person”

    Trước hết, mission này rất phù hợp với beginner: CEFR mô tả A1 có thể hỏi/đáp các câu đơn giản về thông tin cá nhân; Companion Volume thậm chí đặt việc nói tên và hỏi tên người khác ở vùng Pre-A1/A1.
    Portal
    +1

    Finding 1 — n=5 chỉ nên chứng minh pipeline feasibility, không phải learning efficacy

    CONSORT guidance cho pilot/feasibility nhấn mạnh objective phải là feasibility, có progression criteria định trước, và không nên làm formal hypothesis testing efficacy vì sample pilot thường underpowered.
    DOI
    +1

    Confidence: High cho nguyên tắc methodology; indirect vì guidance này không riêng language-learning.

    Recommendation: ADOPT.

    5 người của bạn nên trả lời các câu:

    Có capture đúng attempt không? Có phân biệt unsupported/supported không? Có state nào promote sai? Delayed job có chạy đúng lag? Transfer item có thật sự held-out? Scorer có consistency không?

    Không hỏi: “FlashDay tăng English bao nhiêu %?”

    Finding 2 — đừng train quá nhiều capability cùng lúc trong pilot đầu

    Literature không có bằng chứng kiểu “A1 nên học chính xác 2 capabilities/session”. Vì thế nếu ai đưa 3, 5 hay 7 như một con số khoa học thì đó là false precision.

    Tuy nhiên task repetition có lợi cho L2 oral performance, đặc biệt accuracy; exact repetition thường giúp acquisition ban đầu, trong khi transfer cần sau đó thay đổi task/context có kiểm soát.
    ScienceDirect
    +1

    Confidence: Medium cho sequencing; Low cho con số 2.

    Recommendation: TEST — khởi đầu với 2 target capabilities.

    Mission đầu tiên tôi sẽ dùng:

    Văn bản thuần túy
    meet_new_person
    ├── greet_person [prerequisite / lightweight]
    ├── give_name [target capability]
    └── ask_name [target capability]

    Không nên cùng lúc đo greeting + ask_name + give_name + where_from + age + occupation + closing. Nếu learner fail, causal attribution của evidence engine sẽ rất kém.

    Buổi pilot tôi sẽ triển khai khoảng như sau

    Baseline — 2–3 phút

    Learner thấy tình huống:

    You meet someone for the first time.

    Không exemplar, không hint.

    Hệ thống tạo independent baseline cho give_name và ask_name.

    Quan trọng: baseline không được dạy ngầm trước khi đo.

    Input/model — khoảng 3 phút

    Một micro-dialogue hoàn chỉnh, ví dụ:

    Hi. I'm Minh.
    Hi, I'm Anna.
    What's your name?

    Sau đó một exemplar thứ hai với wording/context hơi khác.

    Đây mới là EXPOSED.

    Supported retrieval — khoảng 4–5 phút

    Cue → learner attempts → feedback → retry.

    Corrective feedback nói chung có tác động bền trong SLA; classroom meta-analysis cũng thấy prompts thường cho kết quả tốt hơn recasts, nhất là ở constructed responses.
    Wiley Online Library
    +1

    Do đó:

    Văn bản thuần túy
    wrong/blank
    → explicit model + short explanation
    → retry

    partial but communicative
    → elicitation/prompt
    → self-repair

    Đừng auto-correct mọi partial response thành câu hoàn chỉnh ngay.

    Independent retrieval — khoảng 4–5 phút

    Capability A → capability B → quay lại A.

    Mục tiêu là retrieval thật, không phải imitation ngay sau model.

    Near-variation probe — khoảng 2 phút

    Thay interlocutor hoặc preceding turn, nhưng giữ underlying communicative goal.

    +24h

    Một no-hint retention probe.

    Sau khi pass retention

    Một held-out transfer probe.

    Minimum evidence package / capability

    Đây là phần tôi nghĩ FlashDay nên làm nghiêm hơn hầu hết learning apps.

    Tôi sẽ không cho state engine promote chỉ từ một score: correct. Một capability tối thiểu cần evidence lineage như sau:

    Evidence Cần để claim
    baseline_attempt raw answer/audio, no scaffold, context ID
    exposure exemplar/version learner thực sự thấy
    retrieval_attempt[] raw response, timestamp, cue, scaffold-before-response
    feedback loại feedback + target error + linked attempt
    retry learner-generated response sau feedback
    independent_confirmations ≥2 no-scaffold successes theo rule pilot
    delayed_probe actual lag, no scaffold, no pre-answer feedback
    transfer_probe context delta + dimension changed + held-out flag
    transition_record state_before, state_after, evidence IDs, rule version

    Ngoài ra scorer nên lưu riêng:

    Văn bản thuần túy
    communicative_success
    form_accuracy
    intelligibility
    scaffold_level
    repair_required
    latency

    Đừng biến chúng thành một opaque 0.82.

    Ví dụ A1 nói:

    “Your name?”

    Đó không nên là FAIL giống hệt im lặng hoặc nói câu không liên quan. Learner đã truyền đạt đúng communicative intent dù form chưa đạt target.

    Confidence: Medium. Đây là system-design inference dựa trên literature retrieval/feedback, không phải một standard research package đã tồn tại.

    Recommendation: ADOPT, nhưng xem threshold promotion là versioned experimental policy.

    1. TRANSFER DIMENSIONS — cái gì thật sự là context change?

    Meta-analysis lớn về transfer của retrieval practice tổng hợp 192 effect sizes / 122 experiments / N=10,382 cho thấy retrieval practice có transfer benefit trung bình d≈0.40, nhưng mức transfer phụ thuộc mạnh vào similarity, initial retrieval success, elaboration và loại transfer.
    PubMed
    +1

    Một nguyên tắc quan trọng trong transfer literature là:

    surface có thể đổi, nhưng deep structure / function cần giữ nguyên nếu bạn đang test transfer của cùng một capability.

    Người mới đặc biệt dễ bám surface cues thay vì nhận ra underlying structure.
    Taylor & Francis Online
    +1

    Với:

    Văn bản thuần túy
    capability = ask_name
    function = obtain interlocutor's name
    canonical realization = "What's your name?"

    tôi sẽ classify 4 dimension của bạn như sau.

    Partner — meaningful nếu đổi cả conversational cues

    Training:

    Anna: Hi, I'm Anna.
    Learner: What's your name?

    Transfer:

    New person: Nice to meet you. I'm new here.
    Learner: ...

    Đây là good near transfer.

    Nhưng:

    Văn bản thuần túy
    avatar Anna → avatar David
    mọi text/cue khác y hệt

    chỉ là surface decoration.

    Confidence: Medium-high.

    Recommendation: ADOPT làm transfer dimension đầu tiên.

    Setting — meaningful khi setting thay đổi discourse demands

    Ví dụ:

    Training: gặp bạn mới ở lớp.

    Transfer: gặp người mới tại một event.

    Nếu background hình ảnh đổi nhưng prompt vẫn nguyên văn:

    “Ask the person's name.”

    thì learner vẫn đang respond vào system cue, không phải context.

    Một context meaningful phải làm cue khác đi:

    Văn bản thuần túy
    training:
    "Hi, I'm Lan."

    transfer:
    "You sit next to someone new before class.
    They smile and say hello."

    Learner phải infer action.

    Confidence: Medium.

    Recommendation: ADOPT, sau partner/lead-in.

    Topic — thường KHÔNG đủ cho ask_name

    “football conversation” → “music conversation”, rồi app vẫn hỏi:

    Ask their name.

    Deep action không đổi và system instruction vẫn giveaway answer.

    Đây phần lớn là surface variation.

    Nếu topic làm thay đổi pragmatic function:

    “What should I call you?”
    “May I have your full name for the booking?”

    thì vấn đề ngược lại: bạn có thể đã chuyển sang một capability/pragmatic register khác, không còn clean test của casual ask_name.

    Confidence: Medium.

    Recommendation: REJECT topic-only như proof of transfer cho mission đầu tiên.

    Channel — meaningful nhưng quá mạnh cho transfer đầu tiên

    Training bằng typed text rồi transfer sang listening + spoken output thay đổi:

    input modality,
    decoding requirement,
    production modality,
    pronunciation,
    temporal pressure.

    Nếu learner fail, bạn không biết ask_name knowledge không transfer hay họ chỉ chưa nghe/ nói được.

    CEFR cũng tách spoken interaction khỏi written interaction.
    rm.coe.int
    +1

    Confidence: High rằng đây là confounded construct change.

    Recommendation: ADAPT.

    Đừng dùng text → voice làm first transfer gate. Hãy coi nó như later cross-channel transfer.

    Một finding rất đáng chú ý: đừng variation quá sớm

    Controlled research cho thấy contextual diversity thường giúp representations/generalisation, kể cả word learning; một study lớn tìm thấy learning ở varied contexts giúp performance tốt hơn trên unfamiliar contexts.
    PubMed Central (PMC)
    +1

    Nhưng variation không phải lúc nào càng sớm càng tốt. Một experiment về retrieval + context variation cho thấy retrieval trong context mới có thể giảm recall, và hiệu ứng xấu này còn carry sang final test.
    PubMed

    Điều này rất quan trọng cho absolute beginner.

    Confidence: Medium-high.

    Recommendation: ADOPT “anchor → vary”, không phải “vary from trial #1”.

    Tức là:

    Văn bản thuần túy

    1. stable exemplar
    2. successful retrieval
    3. successful retrieval again after interference
    4. then near-context variation
    5. delayed retention
    6. held-out transfer
      Ordering cho multi-mission curriculum

    Tôi sẽ không order bằng “học 10 câu rồi sang mission khác”. Tôi sẽ tăng transfer distance dần:

    Văn bản thuần túy
    Stage 1 — SAME STRUCTURE
    same channel
    same intent
    same register
    small surface variation

    Stage 2 — NEAR TRANSFER
    new partner
    new conversational lead-in

    Stage 3 — PRAGMATIC VARIATION
    new setting
    slightly different conversational sequence

    Stage 4 — COMPOSITION
    greet → give_name → ask_name
    without explicit per-capability cue

    Stage 5 — MISSION TRANSFER
    ask_name appears inside another mission
    e.g. "first day at class" / "join a club"

    Stage 6 — CROSS-CHANNEL
    text → listening/speaking
    or voice → live-style interaction

    Task repetition literature supports using repetition first for accuracy/structural control, while transfer literature supports later testing across changed demands rather than assuming exact-task mastery generalizes automatically.
    ScienceDirect
    +1

    Recommendation: ADOPT.

    1. DOSAGE — retrieval, spacing, delayed test, remediation

    Đây là phần cần tránh nhất việc tạo “magic threshold”.

    Retrieval practice: bằng chứng mạnh, nhưng quantity không tuyến tính

    Meta-analyses nói retrieval/practice testing nhìn chung tốt hơn restudy.
    Sage Journals
    +1

    Trong L2 cụ thể, Nakata cho 98 Japanese learners học English–Japanese word pairs với 1 / 3 / 5 / 7 retrievals. 5 và 7 cho final scores cao hơn 1 và 3; nhưng khi kiểm soát time-on-task, 1 retrieval có gain/time lớn nhất.
    DOI

    Điều đó chống lại policy:

    Văn bản thuần túy
    "càng nhiều retrieval trong cùng buổi càng tốt"

    Confidence: High cho vocabulary learning; Medium khi extrapolate sang spoken capability.

    Recommendation: REJECT fixed 5–7 retrieval rule.

    Spacing: đây là phần có evidence L2 rất mạnh

    Meta-analysis Kim & Webb:

    48 experiments,
    N=3,411,
    98 effect sizes,
    spacing cho medium-to-large benefit,
    longer spacing tốt hơn shorter spacing trên delayed tests,
    equal vs expanding spacing không khác rõ.
    Wiley Online Library
    +1

    Một reported aggregate cho delayed effects của spacing đạt khoảng g=0.80; comparison longer vs shorter spacing ở delayed tests khoảng g=0.40.
    ResearchGate
    +1

    Confidence: High.

    Recommendation: ADOPT.

    FlashDay không nên cố đạt mastery bằng 8 lượt liên tiếp. Hãy đưa learner tới mức có thể retrieve, rồi để forgetting xảy ra một chút.

    Successive relearning quan trọng hơn mass repetition

    Rawson & Dunlosky review successive relearning: learner retrieve tới criterion, sau đó relearn ở những sessions cách nhau. Trong experimental work, một relearning session đã tăng retention mạnh, thêm sessions tiếp tục tăng, nhưng lợi ích biên giảm; họ cảnh báo không có một con số “3 sessions là đủ” áp dụng cho mọi context.
    Sage Journals
    +1

    Confidence: High cho memory learning; Medium cho FlashDay oral interaction.

    Recommendation: ADOPT principle, TEST exact schedule.

    Threshold tôi sẽ ship trong pilot

    Không gọi các số sau là “science threshold”. Tôi sẽ version chúng:

    Văn bản thuần túy
    evidence_policy_version = "pilot-v1"
    SUPPORTED → INDEPENDENT

    Require:

    Văn bản thuần túy
    2 no-scaffold successful retrievals

    với điều kiện:

    không consecutive imitation,
    có capability khác hoặc distractor chen giữa,
    feedback của attempt trước không còn visible,
    ít nhất một trial yêu cầu learner tự generate.

    Tại sao không 1?

    Một correct response có thể do residual working memory/cueing.

    Tại sao không 5?

    Dễ biến acquisition thành overtraining và làm pilot dài; evidence không chứng minh 5 là universal optimum.

    Recommendation: TEST.

    Eligibility cho delayed test

    Sau khi learner đã đạt INDEPENDENT, dừng. Đừng bắt họ perform thêm 5 lần “cho chắc”.

    Schedule +24h.

    24h là một delayed assessment thực sự chứ không immediate posttest; meta-analysis L2 phân biệt delayed posttests và cho thấy spacing effects rõ ở delayed outcomes.
    Wiley Online Library

    Nhưng tôi sẽ đổi semantics từ:

    Văn bản thuần túy
    RETAINED

    thành internally:

    Văn bản thuần túy
    RETAINED {
    horizon: "24h"
    }

    hoặc:

    Văn bản thuần túy
    retention_evidence:
    lag_hours: 24.7

    Vì “retained” không nên ngầm hiểu là nhớ lâu dài.

    Confidence: High.

    Recommendation: ADAPT state semantics.

    Bao nhiêu context variation trước delayed test?

    Cho pilot A1:

    1 near-variation exposure/retrieval là đủ trước delayed test.

    Quan trọng hơn quantity là:

    giữ lại ít nhất một unseen variation làm held-out transfer test.

    Nếu bạn luyện:

    Anna,
    David,
    café,
    classroom,
    voice,
    text,

    rồi final test dùng một trong số đó, bạn đã ăn mất transfer set.

    Tôi muốn:

    Văn bản thuần túy
    TRAIN
    context A
    context B = near variation

    24H RETENTION
    context structurally matched to training
    no aid

    TRANSFER
    context C = held-out dimension

    Confidence: Medium.

    Recommendation: TEST.

    Fail vs partial: remediation trigger

    Đây là chỗ FlashDay không nên dùng binary correctness.

    Tôi sẽ dùng ba outcomes.

    PASS

    Learner tự tạo response và đạt communicative goal:

    “What's your name?”

    Hoặc một acceptable realization:

    “What is your name?”

    Không scaffold.

    → count successful retrieval.

    PARTIAL

    Communicative intent vẫn recoverable:

    “Your name?”

    hoặc minor morphology/pronunciation error nhưng interlocutor vẫn hiểu ngay.

    → không reset capability.

    Cho prompt:

    “Try again as a complete question.”

    Learner phải self-repair.

    Oral CF meta-analysis cho thấy corrective feedback có durable effects và prompts thường vượt recasts.
    ERIC

    → PARTIAL không count là full independent success, nhưng cũng không phải fail.

    FAIL
    blank,
    wrong intent,
    unusable answer,
    needs translation/full model,
    learner cannot repair after light prompt.

    → explicit feedback/model → retrieval retry.

    Nếu:

    Văn bản thuần túy
    2 FAILs on same capability inside a short window

    tôi sẽ trigger remediation, không tiếp tục brute-force testing.

    Ví dụ:

    Văn bản thuần túy
    FAIL
    → cue / model
    → retry
    → FAIL
    → decompose:
    comprehension check
    form recognition
    controlled production
    → rebuild retrieval

    2 FAIL là product rule để test, không literature constant.

    Recommendation: TEST.

    State machine tôi sẽ dùng cho FlashDay pilot
    Văn bản thuần túy
    NOT_SEEN
    │
    │ independent baseline
    ▼
    EXPOSED
    │ input/model encountered
    ▼
    SUPPORTED
    │
    │ successful production with help
    │ + corrective feedback / repair
    ▼
    INDEPENDENT
    │
    │ 2 verified no-scaffold successes
    │ separated by intervening retrieval
    ▼
    RETAINED
    │
    │ no-scaffold pass after ≥24h
    │ retention_horizon recorded
    ▼
    TRANSFERRED
    │
    │ no-scaffold success
    │ in HELD-OUT context
    │ preserving communicative function
    ▼
    [future evidence]

    Tôi không cho phép:

    Văn bản thuần túy
    SUPPORTED → RETAINED

    hay:

    Văn bản thuần túy
    one immediate correct answer → INDEPENDENT

    và đặc biệt không cho:

    Văn bản thuần túy
    trained variation → TRANSFERRED
    Recommendation matrix
    Finding Evidence strength FlashDay
    Retrieval > restudy nói chung High — meta-analyses ADOPT
    Feedback đi cùng retrieval rất quan trọng High — meta + SLA studies ADOPT
    Spacing tốt hơn massing cho L2 retention High — 48 experiments ADOPT
    More same-session retrieval luôn tốt hơn Không được support REJECT
    24h đủ chứng minh durable retention Không REJECT
    24h là first retention checkpoint Strongly defensible ADOPT
    Varied contexts hỗ trợ generalisation Medium–High ADOPT
    Variation ngay từ retrieval đầu của beginner Mixed evidence REJECT/ADAPT
    Partner/lead-in change cho ask_name Medium–High inference ADOPT
    Topic-word change = transfer Low REJECT
    text→voice là clean same-capability transfer Low / confounded REJECT initially
    2 unaided successes → Independent Product hypothesis TEST
    2 fails → remediation Product hypothesis TEST
    2 target capabilities/session Product/instrumentation hypothesis TEST

    Nếu phải chốt một thay đổi quan trọng nhất cho evidence engine hiện tại: hãy biến TRANSFERRED từ “pass một task khác” thành “pass một held-out context whose delta is explicitly encoded while communicative deep structure remains constant.” Đây là khác biệt giữa một state machine chỉ ghi completion và một engine thực sự đo evidence of learning.
    PubMed
    +1

  3. Thunderkill016 commented on Sep 30, 2026

    @Thunderkill016
    OwnerAuthor

    Research Round 2 — pilot harness design (ChatGPT, Playwright)

    Hỏi: session structure (day0/+24h/+48h + có nên thêm +72h), claim criteria, unresolved-remediation có block claim không, stress archetypes, cohort metrics không-vanity.

    Quyết định đã encode vào src/vnext/pilot-harness.js (issue #50)

    Finding Decision Encoding
    TRANSFERRED không được là input của chính claim (circular) — recompute từ primitives ADOPT evaluateClaim đếm trên verified primitive events; milestone chỉ là output — mismatch → integrityFailures
    Baseline attribution — pass baseline → không phải của FlashDay ADOPT acquisitionSource: 'PREEXISTING'; loại khỏi claimsEligible denominator
    Rule "no fail after last independent" quá cứng REJECT literal → ADOPT unresolved-contradiction openProbes: latest outcome của từng probe type (delayed/transfer/checkpoint) phải là success; fail→remediate→re-pass = resolved
    +72h probe riêng, không phải claim condition ADOPT retained72h report field riêng, không vào learnedByFlashday
    Oscillator archetype (boss fight) ADOPT learner.osc trong tests/vnext-pilot.test.mjs
    Cohort metrics: claim-complete rate / eligible + 5 metrics phụ ADOPT summary: claimRate (eligible-only), baselineMastered, needsRelearning, integrityFailures
    Vanity metrics (XP/attempts/score avg) REJECT as success criteria không có trong summary
    Harness gate: zero invalid promotions, reconstructable claims, bounded failure paths ADOPT integrityFailures === [] là assert trong test; forged-event case trong §10

    Failed assessment phải re-probe được

    Điểm kỹ thuật nảy ra khi build oscillator: assessment gate trong nextMissionTask cũ chỉ serve khi task chưa attempt → fail là dead-end vĩnh viễn. Fix: re-serve khi latest checkpoint outcome ≠ success, vẫn gate trên TRANSFERRED (src/vnext/mission-runner.js — assessment re-probe). Retry vẫn fresh như lần đầu.

    Full response (raw)

    ChatGPT R2 response
  4. Thunderkill016 commented on Sep 30, 2026

    @Thunderkill016
    OwnerAuthor

    #49 (comment)

    1. SESSION STRUCTURE

    Recommendation: giữ day0 teach → +24h delayed + transfer → +48h fresh assessment, và thêm một probe +72h trong synthetic harness.

    Evidence strength:

    HIGH: spacing và retention phụ thuộc retention interval; L2 meta-analysis 48 experiments, N=3,411 cho thấy spacing có lợi và longer spacing đặc biệt có lợi trên delayed tests.
    DOI
    +1
    HIGH: successive relearning nên diễn ra qua nhiều spaced sessions thay vì massed practice trong một buổi.
    Hiệp hội Khoa học Tâm lý
    LOW cho chính xác 72h: literature không cho thấy 72h là ngưỡng đặc biệt.

    Verdict: ADOPT cho synthetic harness / TEST cho human pilot.

    Tôi sẽ dùng:

    Văn bản thuần túy
    T0 baseline → teach → independent
    +24h delayed retention
    then held-out transfer
    +48h fresh assessment
    +72h retention-horizon probe

    Nhưng +72h phải là probe riêng, không phải điều kiện bắt buộc để claim TRANSFERRED.

    Lý do: synthetic harness gần như không có participant-cost, nên bỏ +72h là mất một boundary rất đáng test. Nó cho phép engine phân biệt:

    Văn bản thuần túy
    retained_24h = true
    retained_72h = false

    thay vì một boolean mơ hồ RETAINED=true.

    Ngoài ra, với n≈5 synthetic learner, đây vẫn là feasibility/state-machine validation, không phải evidence rằng cadence này tối ưu cho người học thật. Pilot methodology khuyến nghị prespecify feasibility criteria thay vì efficacy claims.
    PubMed Central (PMC)
    +1

    1. CLAIM CRITERIA

    Claim hiện tại của bạn:

    Văn bản thuần túy
    TRANSFERRED

    • ≥2 unaided successes spaced
    • ≥1 delayed ≥24h
    • ≥1 held-out transfer
    • assessment pass

    Recommendation: gần đủ, nhưng sửa 3 điểm.

    A. Đừng dùng TRANSFERRED làm evidence đầu vào của chính claim

    Nếu TRANSFERRED vốn được derive từ delayed + transfer evidence thì:

    Văn bản thuần túy
    TRANSFERRED
    AND delayed
    AND transfer

    là circular/redundant.

    Claim nên recompute từ primitive evidence:

    Văn bản thuần túy
    LEARNED :=
    baseline_did_not_already_demonstrate_capability
    AND independent_successes >= 2
    AND independent_successes_are_spaced
    AND delayed_success(lag >= 24h)
    AND held_out_transfer_success
    AND fresh_assessment_success
    AND evidence_package_complete
    AND no_unresolved_contradiction

    Sau đó:

    Văn bản thuần túy
    milestone = TRANSFERRED

    là output của predicate đó.

    Evidence strength: MEDIUM — inference từ sound measurement/provenance design, không phải learning-science threshold.

    Verdict: ADOPT.

    B. Thêm baseline attribution

    Đây là phần còn thiếu quan trọng nhất.

    Nếu learner:

    Văn bản thuần túy
    baseline ask_name → PASS

    thì FlashDay không được claim learned capability đó.

    Nó nên thành kiểu:

    Văn bản thuần túy
    BASELINE_MASTERED
    // hoặc:
    acquisition_source = PREEXISTING

    Learner vẫn có thể chạy delayed/transfer/assessment để xác nhận capability, nhưng không được tính vào:

    Văn bản thuần túy
    learned_by_flashday

    Evidence strength: HIGH về causal/measurement logic.

    Verdict: ADOPT.

    C. Không yêu cầu “không có fail nào sau independent”

    Rule:

    no attempt nào fail sau lần independent cuối

    là quá cứng.

    Learner thật có thể:

    Văn bản thuần túy
    independent PASS
    24h FAIL
    remediation
    independent PASS
    24h PASS
    transfer PASS
    assessment PASS

    Đây vẫn là acquisition hợp lệ.

    Thay bằng:

    Văn bản thuần túy
    no_unresolved_contradictory_evidence_at_claim_time

    Ví dụ:

    Văn bản thuần túy
    latest relevant delayed = FAIL
    → claim blocked

    FAIL
    → remediation
    → new independent evidence
    → new delayed PASS
    → contradiction resolved

    Không xóa failure cũ; giữ nguyên provenance.

    Evidence strength: MEDIUM cho policy cụ thể; forgetting/relearning qua spaced sessions có support mạnh.
    Hiệp hội Khoa học Tâm lý

    Verdict: REJECT literal rule → ADOPT unresolved-contradiction rule.

    1. FAILURE-PATH LEARNERS

    Trong 3 archetype bạn nêu:

    (b) delayed fail ×2 → remediation → requalification

    Stress engine mạnh nhất về state machine.

    Nó kiểm tra:

    Văn bản thuần túy
    INDEPENDENT
    → delayed FAIL
    → remediation
    → SUPPORTED?
    → regain INDEPENDENT
    → reschedule delayed
    → PASS
    → transfer...

    Các bug dễ lộ:

    infinite remediation loop
    stale delayed event promote learner
    previous PASS thắng newer FAIL
    duplicate retry tạo double evidence
    state không demote/reopen
    scheduled assessment vẫn chạy dù prerequisite invalid
    remediation success tự động thành INDEPENDENT

    Evidence strength: HIGH rằng spaced relearning cần repeated successful retrieval across sessions; chính transition policy là FlashDay-specific.
    Hiệp hội Khoa học Tâm lý

    Verdict: ADOPT bắt buộc.

    (c) ASR/self-report-only

    Stress evidence trust mạnh nhất.

    Self-assessment chỉ tương quan vừa với measured language performance trong meta-analysis 67 studies / >68,500 participants (r≈.466), nên không nên thay performance evidence.
    ERIC

    Automated speech scoring có thể hữu ích, nhưng validation vẫn là vấn đề; một study còn thấy automarker lenient hơn với low-proficiency speakers.
    Taylor & Francis Online

    Nếu chỉ có:

    Văn bản thuần túy
    self_report
    ASR transcript

    và không có auditable raw performance / validated scorer evidence, tôi đồng ý:

    Văn bản thuần túy
    max_state = SUPPORTED

    Evidence strength:

    self-report: HIGH để bác việc dùng nó như objective mastery evidence.
    ASR-only cap: MEDIUM, phụ thuộc scorer validation.

    Verdict: ADOPT fail-closed.

    Sau này nếu bạn lưu:

    Văn bản thuần túy
    raw_audio
    ASR_model_version
    confidence
    pronunciation/intent scorer version
    acceptance thresholds

    và validate scorer, có thể nới rule.

    (a) baseline PASS → skip teaching

    Cũng phải có, nhưng nó stress attribution/short-circuiting hơn learning loop.

    Expected:

    Văn bản thuần túy
    baseline PASS
    → DO NOT teach unnecessarily
    → mark pre-existing
    → schedule confirmation/transfer if desired
    → never emit learned_by_flashday

    Evidence strength: HIGH measurement logic.

    Verdict: ADOPT.

    Tôi sẽ thêm archetype thứ 5: oscillator

    Đây mới là “boss fight”:

    Văn bản thuần túy
    baseline FAIL
    teach
    independent PASS ×2

    +24h FAIL
    remediation
    independent PASS

    delayed PASS
    transfer FAIL
    targeted variation/remediation
    transfer PASS

    assessment FAIL
    re-assess
    assessment PASS

    +72h FAIL

    Expected final state có thể là:

    Văn bản thuần túy
    TRANSFERRED = historical evidence true
    assessment_current = pass
    retained_24h = true
    retained_72h = false
    needs_relearning = true

    Engine không được collapse nó thành đơn giản:

    Văn bản thuần túy
    LEARNED=true

    Evidence strength: MEDIUM as harness design.

    Verdict: ADOPT.

    1. COHORT REPORT

    Với 5 synthetic learners, không lấy average score làm headline.

    Primary metric
    claim-complete capability rate
    Văn bản thuần túy

    capabilities satisfying full claim predicate

    /

    capabilities eligible to be learned

    Exclude BASELINE_MASTERED khỏi denominator của “learned by FlashDay”.

    Ví dụ:

    Văn bản thuần túy
    7 / 10 eligible capabilities
    full-evidence TRANSFERRED

    Evidence strength: HIGH cho feasibility framing. Pilot guidance khuyến nghị outcome gắn trực tiếp vào feasibility/progression criteria.
    BMJ
    +1

    Verdict: ADOPT.

    Tôi sẽ report đúng 6 metric
    Metric Verdict
    Full-evidence claim rate ADOPT — primary
    State distribution after each session (SUPPORTED/INDEPENDENT/...) ADOPT
    Attempts/retrievals per claimed capability ADOPT
    Remediation episodes per capability ADOPT
    24h→72h retention survival TEST
    Evidence integrity failures / invalid promotions ADOPT — critical

    Đặc biệt metric cuối với headless harness phải bằng:

    Văn bản thuần túy
    invalid_claims = 0
    missing_evidence_promotions = 0
    stale_event_promotions = 0

    Nếu TRANSFERRED rate = 100% nhưng có một learner được promote bằng stale/partial evidence thì harness fail, không phải thành công.

    Không dùng làm headline
    Văn bản thuần túy
    total XP
    total attempts
    completion rate
    average score
    time/session
    number of events generated

    Chúng chỉ diagnostic, chưa nói engine đo learning đúng.

    Verdict: REJECT as success metrics.

    Tôi sẽ chốt harness gate như sau
    Văn bản thuần túy
    HARNESS PASS iff

    1. Every positive learning claim is reconstructable
      from primitive evidence.

    2. Zero claims are produced with:

      • missing provenance
      • only self-report
      • unresolved later contradictory evidence
      • trained context mislabeled as transfer
      • baseline-mastered capability attributed to FlashDay.
    3. Every expected failure path terminates:

      • remediation has bounded/reachable exit
      • stale scheduled events cannot promote state
      • retries are idempotent.
    4. Report separately:
      learned_by_flashday
      baseline_mastered
      retained_24h
      retained_72h
      transferred
      current_assessment_status.

    Evidence strength: MEDIUM for exact engineering gate; underlying feasibility/retrieval/spacing rationale HIGH.

    Verdict: ADOPT.

    Điểm quan trọng nhất Round 2: đừng để TRANSFERRED trở thành truth source. Primitive evidence mới là truth; milestone chỉ là materialized conclusion có thể recompute. Đây sẽ giúp FlashDay tránh rất nhiều bug khi delayed failure, remediation và reassessment bắt đầu chồng lên nhau.

  5. Thunderkill016 commented on Sep 30, 2026

    @Thunderkill016
    OwnerAuthor

    Research Round 3 — TEST hypotheses → policy, workstream ordering (ChatGPT)

    Hỏi: pilot cohort xong (3/4 learned, able=PREEXISTING, osc=needsRelearning, integrity=0). TEST hypotheses nên là engine constant hay config? Workstream tiếp theo (UI vs curriculum vs persistence)?

    Verdicts

    Item Verdict Đã encode?
    2 unaided spaced → engine constant REJECT giữ ở claim/policy level
    2 unaided default versioned policy ADOPT sẽ encode vào policy.independent
    2 fails → remediation engine REJECT planner opt consecutiveFailuresBeforeRemediation
    2 caps/session engine REJECT curriculum config, không vào engine
    24h retention threshold hard-code ADAPT policy.retention.minLagMs; lag từ timestamps, không từ session label
    12 engine invariants (immutability, provenance, no absorbing dead-end…) ADOPT hầu hết đã pin trong tests
    Firebase persistence NEXT ADOPT #1 workstream mới
    Minimal honest UI ADOPT #2 sau persistence
    Scale A1 DAG ngay REJECT → ADOPT #3 sau human interaction đầu

    Quote chốt: "Engine defines what happened. Policy defines how much evidence is enough."

    Full response

    ChatGPT R3

    Pilot này cho thấy engine đã qua được bài test quan trọng: historical milestone ≠ current mastery state. Oscillator giữ TRANSFERRED lịch sử nhưng claim bị chặn bởi evidence mới là đúng hướng.

    1. Cái gì hard-code, cái gì config?
      2 unaided spaced successes → INDEPENDENT

    Verdict: ADAPT — policy constant, KHÔNG engine constant.

    Engine nên hard-code:

    Văn bản thuần túy
    attempt.scaffoldLevel
    attempt.success
    attempt.timestamp
    lagFromPreviousSuccess

    và khả năng query:

    Văn bản thuần túy
    countUnaidedSuccessfulRetrievals(...)

    Nhưng:

    Văn bản thuần túy
    requiredIndependentSuccesses = 2
    minSpacing = ...

    phải nằm trong versioned learning policy.

    Lý do: repeated successful retrieval có evidence tốt, nhưng criterion tối ưu thay đổi theo material, learner và việc có subsequent relearning hay không. Các nghiên cứu thậm chí cho thấy lợi ích của criterion ban đầu cao bị “override” khi có successive relearning.
    PubMed
    +1

    Hard-code 2 hôm nay sẽ khiến sau này muốn:

    Văn bản thuần túy
    ask_name -> 2
    pronunciation distinction -> 3
    receptive recognition -> 1

    phải migration semantics toàn engine.

    Risk hard-code sai: HIGH.

    Tôi sẽ để:

    TypeScript
    policy.independent = {
    successfulUnaidedRetrievals: 2,
    requireSpacing: true,
    minLagMs: ...
    }

    và mỗi claim lưu:

    Văn bản thuần túy
    policyVersion
    2 failures → remediation

    Verdict: TEST / PRODUCT-CONFIG.

    Tuyệt đối không hard-code trong projection.

    Engine chỉ nên biết:

    Văn bản thuần túy
    FAIL
    PARTIAL
    SUCCESS

    và chronology.

    Planner quyết định:

    Văn bản thuần túy
    consecutiveFailuresBeforeRemediation = 2

    Vì một lỗi có thể cần remediation ngay:

    Văn bản thuần túy
    learner hoàn toàn không hiểu intent

    trong khi một lỗi pronunciation nhẹ có thể cho thêm retrieval.

    Unsuccessful retrieval bản thân nó cũng không vô ích nếu sau đó có corrective restudy; research về test-potentiated learning cho thấy failed retrieval có thể tăng hiệu quả encoding tiếp theo.
    PubMed

    Risk hard-code sai: HIGH.

    2 capabilities/session

    Verdict: TEST — curriculum/planner config בלבד.

    Không được xuất hiện trong evidence engine.

    Đây là:

    Văn bản thuần túy
    session workload policy

    không phải definition của learning.

    Sau này mission có thể là:

    Văn bản thuần túy
    meet-person:
    ask_name
    give_name

    nhưng mission khác chỉ có 1 capability khó hoặc 4 micro-capability rất nhỏ.

    Risk hard-code sai: VERY HIGH, vì nó leak curriculum assumptions xuống domain engine.

    ≥24h delayed

    Verdict: ADAPT — lag là engine fact, threshold là policy.

    Hard-code:

    Văn bản thuần túy
    actualLagMs

    Không hard-code:

    Văn bản thuần túy
    24h === retention

    Nên:

    Văn bản thuần túy
    policy.retention.minLagHours = 24

    và evidence lưu horizon thật:

    Văn bản thuần túy
    retainedAtHours = [25]
    failedAtHours = [75, 100]

    Spacing/relearning evidence rất mạnh; nhưng không có một 24h universal boundary.
    PubMed
    +1

    Cái GÌ thật sự nên là ENGINE invariant?

    Verdict: ADOPT.

    Tôi chỉ hard-code những thứ mà nếu sai thì truth model sai, không phải pedagogy chưa tối ưu:

    Văn bản thuần túy

    1. Raw evidence immutable / append-only.

    2. Claim phải reconstruct được từ primitive evidence.

    3. Every derived state carries policyVersion.

    4. PREEXISTING cannot become learnedByFlashday
      without later evidence of newly taught capability
      being a distinct construct.

    5. Trained context cannot count as held-out transfer.

    6. Assisted evidence cannot masquerade as unaided.

    7. Actual lag derives from timestamps,
      never from session label "day1".

    8. Historical milestones remain historical.
      New contradictory evidence does not delete them.

    9. Current claim evaluates latest relevant evidence.

    10. Unsupported/forged/stale revisions cannot alter projection.

    11. Projection/replay deterministic.

    12. Failed checkpoint remains re-probeable;
      failure must never create an absorbing dead-end
      unless an explicit terminal policy says so.

    Đây chính là đường phân cách:

    Engine defines what happened. Policy defines how much evidence is enough.

    Các progression criteria trong pilot research cũng nên prespecify nhưng có thể thay đổi khi evidence/feasibility cho thấy cần chỉnh; đừng nhầm threshold thử nghiệm với immutable truth.
    PubMed Central (PMC)
    +1

    1. Workstream tiếp theo

    Tôi chọn:

    C → A → B
    Persistence/events → minimal honest UI → curriculum DAG
    C. Firebase persistence/events — làm trước, nhưng cực minimal

    Verdict: ADOPT FIRST.

    Không build “cloud architecture lớn”. Chỉ persistence layer cho truth model đã có.

    Invariant trước khi làm UI:

    Văn bản thuần túy
    append(event) exactly-once logically
    ↓
    reload
    ↓
    replay
    ↓
    same projection
    ↓
    same claim

    Bắt buộc:

    Văn bản thuần túy
    eventId idempotent
    learnerId
    missionRunId
    capabilityId
    occurredAt
    recordedAt
    event schema version
    policy version where relevant

    append-only
    ordered/replayable
    offline/retry cannot duplicate learning evidence
    client cannot overwrite prior event

    Test quan trọng nhất:

    Văn bản thuần túy
    headless oscillator
    → persist Firebase
    → kill process
    → reload events
    → replay
    → byte/logically equivalent claim

    Nếu cái này chưa pass, đừng nối UI.

    Nếu không, UI bug/network retry có thể tạo:

    Văn bản thuần túy
    2 attempts → thành 4

    và learner tự nhiên đạt INDEPENDENT.

    A. Minimal honest UI

    Verdict: ADOPT SECOND.

    Không design FlashDay “đẹp” lúc này.

    UI đầu tiên chỉ cần thể hiện chính xác:

    Văn bản thuần túy
    INPUT
    → RETRIEVAL
    → COMMIT ANSWER
    → FEEDBACK
    → RETRY
    Invariant lớn nhất

    Feedback không được tồn tại trước khi evidence của retrieval được committed.

    Tức là:

    Văn bản thuần túy
    show prompt

    learner answers
    ↓
    ATTEMPT_COMMITTED
    ↓
    only now reveal feedback
    ↓
    retry = NEW attempt

    Không:

    Văn bản thuần túy
    hint appears
    learner closes hint
    answer counted UNAIDED

    Mỗi aid phải contaminate attempt đó:

    Văn bản thuần túy
    hintViewed=true
    modelViewed=true
    translationViewed=true
    replayCount=...

    và không thể undo contamination bằng cách đóng popup.

    Đây là invariant cực quan trọng.

    UI cũng không được expose engine state như game reward

    Không hiện:

    “Bạn đã TRANSFERRED!”

    chỉ vì projection vừa promote.

    Hiện learner-facing language kiểu:

    “Bạn làm được tình huống này mà không cần trợ giúp.”

    Engine state vẫn internal.

    B. A1 curriculum DAG

    Verdict: ADOPT THIRD, chưa scale ngay.

    Lý do không làm trước UI: synthetic learner đã validate state machine, nhưng chưa validate con người hiểu prompt/task thế nào.

    Nếu bạn build 100 capability ngay bây giờ rồi human UI test cho thấy:

    Văn bản thuần túy
    ask_name

    thực ra phải split thành:

    Văn bản thuần túy
    recognize_name_question
    respond_with_name
    initiate_name_question
    repair_when_not_understood

    thì DAG phải sửa hàng loạt.

    Invariant trước khi mở DAG

    Mỗi node phải định nghĩa được:

    Văn bản thuần túy
    capabilityId

    communicative function

    prerequisites

    acceptable realizations

    baseline task

    teaching evidence

    independent task

    near variation

    HELD-OUT transfer dimension

    assessment task

    scoring contract

    remediation path

    Nếu capability không thể tạo held-out assessment khác training mà vẫn giữ cùng construct:

    REJECT node definition.

    Thứ tự triển khai cụ thể
    Văn bản thuần túy
    NOW
    │
    ├─ C1 Firebase append-only event store
    ├─ C2 persistence/replay parity tests
    ├─ C3 adversarial duplicate/out-of-order tests
    │
    ▼
    A1 one ugly-but-honest mission UI
    ├─ input
    ├─ retrieval
    ├─ feedback
    ├─ retry
    └─ reconnect/reload
    │
    ▼
    A2 chạy 2–5 người thật
    │
    ▼
    B1 revise capability ontology
    ├─ confirm granularity
    ├─ prerequisites
    ├─ transfer dimensions
    └─ scoring contracts
    │
    ▼
    B2 expand A1 curriculum DAG
    Verdict cuối
    Decision Verdict
    2 unaided hard-code engine REJECT
    2 unaided default versioned policy ADOPT
    2 fails → remediation engine constant REJECT
    2 fails planner default TEST
    2 capabilities/session engine REJECT
    2 capabilities/session curriculum config TEST
    Evidence/provenance rules hard-coded ADOPT
    Firebase persistence tiếp theo ADOPT #1
    Minimal evidence-honest UI ADOPT #2
    Scale A1 curriculum ngay REJECT
    Small DAG after first human interaction ADOPT #3

    Và bug assessment dead-end mà pilot vừa tìm được là bằng chứng rất tốt cho kiến trúc này: failure phải là evidence, không được vô tình trở thành terminal state. Threshold thay đổi được; lịch sử evidence và các invariant về provenance thì không.

    Bước kế tiếp nên chốt là

    Định nghĩa policy schema versioned

  6. Thunderkill016 commented on Sep 30, 2026

    @Thunderkill016
    OwnerAuthor

    GPT Round 4 (ChatGPT via browser) — Minimal honest UI contract

    $(cat /tmp/r4-post.md)

  7. Thunderkill016 commented on Sep 30, 2026

    @Thunderkill016
    OwnerAuthor

    R5 — curriculum DAG (workstream B, issue #57)

    Full GPT response captured. Decisions encoded below; raw notes available in session.

    Schema split (locked):

    CAPABILITY DAG = dependency of abilities
    MISSION        = composition/context
    POLICY         = sequencing + dosage
    ENGINE         = evidence truth
    

    DAG shape — ADOPT function-granularity: capability = one communicative function that can pass/fail independently (interact.ask_name, never introduce_self). Namespace: reception.listen.*, reception.read.*, production.speak.*, production.write.*, interaction.* — modality in identity only when the construct/evaluator truly differs. Edges only for real ability dependencies; REJECT linear listening→reading→speaking chains and mission-order edges (M1→M2 is sequencing, not a DAG edge).

    Missions — ADOPT: preferred 2 target caps (softMax 3 via policy), 4–6 carriers allowed. Carriers may add retention evidence but can never silently become acquisition teaching or held-out transfer. 7 missions = honest "A1 spoken-communication core coverage X/Y" — NOT "learner reached A1" (that needs a declared CEFR descriptor matrix, ~12–18 mission families).

    Transfer/families — ADOPT contextSignature: {communicativeFunction, interlocutorRole, relationship, setting, register, channel, cueTopology, responseTopology, lexicalDomain}. Canonical family id: pf.<capability>.<cueTopology>.<settingClass>.<register>.<channel>.vN with #variant suffixes inside a family. Same family ⇔ same cue structure + same solution template without context reinterpretation (name-swaps are same-family; open_social vs self_intro topology are different families). Freshness invariant: transfer/fresh familyId ∉ prior practiced familyIds — plus a curriculum linter warning on different-familyId/same-signature fake novelty.

    Layering — locked: ability deps → DAG/content; sequencing/dosage (targets-per-mission, carrier frequency) → versioned policy; invariants (practiced≠transfer, assisted≠unaided, acyclic, valid refs) → engine schema.

    Vietnamese risk — ADAPT: riskTags (e.g. vn.final_consonant, vn.article_omission) drive probes/scoring-attention/feedback/remediation families — never prerequisite edges. "Can I have coffee?" achieves the communicative purpose despite article gaps; linguistic accuracy is not a fake blocker.

    Assessment — ADOPT strict: every claim-bearing target cap needs unaided + delayed + held-out transfer + fresh assessment. Cross-mission sampling is a maintenance layer (refresh retention, detect regression) — never fills a capability's missing evidence package.

    7-mission seed skeleton: M1 meet_new_person (give_name, ask_name) → M2 share_basic_details (state_origin, ask_origin) → M3 meet_at_a_time (understand_time←recognize_numbers, state_time) → M4 order_food_drink (request_item, respond_to_simple_choice) → M5 buy_small_item (ask_price, understand_price←numbers) → M6 find_a_place (ask_location, follow_short_direction←location_terms) → M7 talk_about_self_family (describe_self_basic, describe_person_basic). Carriers: greet/ask_name/give_name reused across later missions.

    Three automated curriculum checks to ship with the schema: (1) DAG validity — acyclic + no dangling prereqs; (2) evidenceability — every target cap has practiced+transfer+fresh families + real evaluator contract; (3) novelty integrity — transfer/fresh families don't alias practiced signatures. If a capability can't satisfy (2), it doesn't enter the graph.

  8. Thunderkill016 commented on Sep 30, 2026

    @Thunderkill016
    OwnerAuthor

    Research R6 — ChatGPT verdict on the 7-mission seed skeleton

    Round 6 validated the proposed M3–M7 skeleton against the schema that
    landed in PR #58. Verdict: direction APPROVED, with three structural
    adjustments (encoded below). Full response at the bottom.

    Decisions encoded (implemented on devin/vnext-curric-001, commit 27f98da)

    1. Prompt-family ids are now injective: the signature hash
      pf.<cap>.<cue>.<setting>.<register>.<channel>.<sigHash8>.vN covers
      the WHOLE contextSignature — R6 flagged that vN was mixing
      "authoring revision" with "semantic family distinction". Fixtures
      derive ids via canonicalFamilyId(cap, sig) — id≡sig by construction.
    2. Role-aware introductions in the planner: targets get a mandatory
      baseline probe; carriers rehearse opportunistically (no baseline —
      that is the point of the role); supports are demand-driven only.
      Mission A dropped four carrier diagnostics + ask_repeat; the drink
      offer check became post-input retrieval.
    3. Surface cap ≤6 declared capabilities/mission (lint in the gate).
    4. interaction.order_drink → interaction.request_item — same
      communicative function; drink vs food is context, not capability.
    5. recognize_numbers_basic is NOT a hard DAG prerequisite — it is
      a support substrate (reception.listen.identify_spoken_number);
      hard edges only after human evidence says the planner needs them.

    M3–M7 authoring table (from R6)

    Mission Targets Carriers Supports
    M3 meet_at_a_time reception.listen.understand_clock_time, production.speak.state_clock_time greet, ask_name identify_spoken_number
    M4 complete_small_order interaction.request_item, interaction.answer_simple_choice greet, thank understand_simple_choice
    M5 buy_small_item interaction.ask_price, reception.listen.understand_spoken_price request_item, thank identify_spoken_number
    M6 find_a_place interaction.ask_location, reception.listen.follow_short_direction greet, thank identify_basic_direction_term
    M7 talk_about_self_family production.speak.state_basic_self_detail, production.speak.describe_family_member_basic say_own_name lexical only

    Transfer signatures per target were specified concretely (e.g.
    understand_clock_time → appointment_confirmation/clinic/receptionist;
    ask_location → locate_destination_inside_building/shopping_center/staff).


    Full ChatGPT response (R6)

    $(cat /tmp/r6-answer.txt)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agent:codexClaimed by Codex agent

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions