Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion biome.json
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,9 @@
"!**/pnpm-lock.yaml",
"!**/package-lock.json",
"!**/yarn.lock",
"!**/docs/supabase-email-templates"
"!**/docs/supabase-email-templates",
"!**/src/vnext",
"!**/src/core"
]
},
"formatter": {
Expand Down
126 changes: 126 additions & 0 deletions docs/flashday/BASELINE_AUDIT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,126 @@
# BASELINE_AUDIT.md — EchoType → FlashDay Next

Baseline verified on the fork `Thunderkill016/flashday-next`,
branch `flashday/evidence-kernel-bootstrap`.

## Provenance

| Item | Value |
|---|---|
| Upstream repo | https://github.com/Talljack/echo-type |
| Fork | https://github.com/Thunderkill016/flashday-next (real GitHub fork, parent link intact) |
| Upstream baseline SHA | `deccdf5e644d5f0b5e2b010f0ac81ee17f66e065` — `feat: 音频边听边校对并直接开始练习 (#121)` |
| Fork remote | `origin` → flashday-next; `upstream` → Talljack/echo-type |
| License | MIT, `LICENSE` unchanged, copyright Talljack retained |

## Baseline gate results (on upstream HEAD, zero local changes)

| Gate | Command | Result |
|---|---|---|
| Install | `pnpm install --frozen-lockfile` | clean, 13s |
| Typecheck | `pnpm typecheck` (`tsc --noEmit`) | PASS, 0 errors |
| Lint | `pnpm lint` (biome, `--diagnostic-level=error`) | PASS, 647 files |
| Unit tests | `pnpm test` (vitest) | **1081 pass / 3 skip**, 185 files, 22s |
| Build | `pnpm build` (Next 16.3.4 Turbopack) | PASS, 69 static pages + full API surface |

Environment: Node v24.16.0, pnpm 9.15.1 (corepack).
E2E (`e2e/*.spec.ts`, Playwright) not run at baseline — requires dev
server; deferred to first slice verification.

## Major subsystems (verified in code, not just docs)

### App shell
- Next.js 16 App Router + React 19 + TypeScript, `(app)` route group.
- Modules: `/listen`, `/read`, `/speak`, `/write`, `/library`,
`/review/today`, `/dashboard`, `/favorites`, `/learn`,
`/pronunciation`, `/weak-spots`, `/journal`, `/settings`.
- API surface: 30+ routes — chat, AI generate, assessment,
pronunciation, recommendations, translate (paid + free), speak, STT,
multi-vendor TTS (Kokoro, Fish, Google, Edge, OpenAI), import
(url/youtube/pdf/transcribe/extract-text), tools (classify, extract,
download), journal OCR, auth callback/token, ollama warmup.

### Persistence (verified)
- Dexie.js IndexedDB, **schema v20** (`src/lib/db.ts`). Tables:
contents, records, sessions, books, conversations, favorites,
favoriteFolders, lookupHistory, translationCache, mediaBlobs,
alignmentCache, collections, weakSpots, journals, learningUnits,
lessons, pronunciationProgress, learningAttempts, dailyTasks,
importJobs, importDrafts, syncConflicts, syncEntityState.
- Supabase: auth + cloud sync. `src/lib/sync/engine.ts` pushes/pulls
12 tables: contents, records, sessions, favorites, favoriteFolders,
journals, books, collections, weakSpots, pronunciationProgress,
learningAttempts, dailyTasks.

### Learning surface (what exists today)
- `LearningRecord` per content item: attempts, accuracy, mistakes,
`fsrsCard` — memory scheduling only.
- `src/lib/fsrs.ts` — thin, clean `ts-fsrs` wrapper; serializable
`FSRSCardData`, `accuracyToRating`, `gradeCard`, `previewRatings`,
`migrateToFSRS`. Keepable as-is — it is the memory model, nothing more.
- `weakSpots` — recurring-error ledger
(`[module+weakSpotType+normalizedText]`, count, resolved). Natural
recurring-error evidence source.
- `LearningAttempt` — immutable submission records ("revisions append,
never replace") with `feedback.source: self|ai`, `usedTranslation`,
`cycle` evidence — already evidence-flavored.
- `text-learning-cycle.ts` — understand → output → correct → recall →
apply stage machine with 24h delay, assisted flags, and a primitive
transfer check (`transferError`: expression must appear in a NEW
context). Derives "practice evidence, not proficiency" — the comment
already knows the invariant.
- `daily-task-planner.ts` — budget-based daily task selection with
evidence application. Closest existing thing to "Today".
- Assessment (`/api/assessment`) — AI-generated adaptive MCQ
(vocab/grammar/reading) → CEFR estimate. Per mission: epistemic role
is **placement estimate**, never verified proficiency.

### Speech / pronunciation
- Web Speech API (recognition + synthesis) + server STT fallback +
Tauri path. Read module: Levenshtein word-level coloring
(green/yellow/red). `/api/pronunciation` = AI-judged eval.
- **Gap for FlashDay**: STT transcript match is treated close to
pronunciation success in places — mission §7 boundary applies:
transcript match ≠ pronunciation evidence.

### State / platform
- 33 Zustand stores (`src/stores/`), hydrated in `(app)/layout.tsx`.
- Tauri v2 desktop (`src-tauri/`), iOS dir present, sidecar Next.js.
- AI: Vercel AI SDK 6, `provider-resolver.ts` with capability
detection, 15+ providers, shared Groq fallback, Upstash rate limit.

## Licenses

| Item | License | Note |
|---|---|---|
| echo-type repo | MIT | retain `LICENSE` + attribution |
| ts-fsrs | MIT | FSRS scheduling lib |
| youtube-transcript | MIT | transcript fetch |
| pdf-parse / pdfjs-dist / mammoth | MIT / Apache-2.0 / BSD-2 | import pipeline |
| assets | bundled | `library-*.png`, icons — repo-internal |

## Major technical risks

1. **Two persistence systems**: Dexie (local-first) + Supabase sync vs
FlashDay's append-only Firestore event log. Bridge must add a Dexie
event table + sync mapping; do NOT dual-write semantics.
2. **Language split**: kernel is plain ESM `.js` (≈12k LOC), shell is
strict TS. Decide vendor-as-is vs port — vendoring preserves the
tested kernel byte-for-byte.
3. **Fire-and-forget AI authority**: several EchoType surfaces treat
AI/STT output as ground truth (pronunciation, assessment). The
bridge must demote them to `ai_llm`/`asr` evaluation authority —
the kernel already refuses independent-ability credit for those.
4. **Angular-seam drift**: EchoType records accuracy per *content*,
kernel needs per *capability/task contract*. Content↔capability
binding is the main design work.
5. **Tauri/desktop**: keep working; all bridge code must be
platform-agnostic (no Node-only deps in the kernel path).
6. **Supabase schema** for a new events table — append-only semantics
need enforcement (RLS or convention); Dexie side is ours to design.

## Out of audit scope (deferred)

- e2e suite state, Tauri build, iOS build
- Real Supabase project config (env-only)
- AI provider key configuration (user-supplied at runtime)
182 changes: 182 additions & 0 deletions docs/flashday/FLASHDAY_ARCHITECTURE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,182 @@
# FLASHDAY_ARCHITECTURE.md — EchoType body, FlashDay brain

> Doctrine: `activity ≠ learning ≠ retention ≠ transfer ≠ proficiency`.
> A score is an observation; a capability claim comes only from
> FlashDay contracts + evidence rules.

## Layers

```
┌─────────────────────────────────────────────────────────┐
│ PRODUCT SHELL (EchoType — keep) │
│ Next.js 16 UI: Listen · Read · Speak · Write · Library │
│ Import (YT/PDF/URL/audio) · TTS/STT transport · Auth │
│ Supabase sync · Tauri desktop · AI provider abstraction │
└──────────────┬──────────────────────────────────────────┘
│ observed activity only (never semantics)
▼
┌─────────────────────────────────────────────────────────┐
│ EVIDENCE BRIDGE (new — the seam this milestone proves) │
│ exercise result ──► registered Task contract │
│ ──► evaluator (authority-stamped) │
│ ──► bindAttempt → EvidenceEvent │
│ UI may supply ONLY observed reality. Semantic fields │
│ (capabilityId, purpose, freshness, outcome credit) are │
│ contract-derived; supplying them throws. │
└──────────────┬──────────────────────────────────────────┘
│ append-only events
▼
┌─────────────────────────────────────────────────────────┐
│ LEARNING KERNEL (FlashDay vNext — vendored, byte-stable) │
│ capabilities (DAG+conditions+evidence reqs) │
│ contracts (Mission/Task/Assessment validators) │
│ evidence (validated append-only events, support flags) │
│ bind (sole mint; forgery-proof) │
│ projection → learner state (separate dimensions) │
│ learner-model (read view, no mastery score) │
│ planner + next-for-you (auditable decisions) │
│ correction-episodes (repair) · support-demand routing │
└──────────────┬──────────────────────────────────────────┘
│ decision + reason
▼
┌─────────────────────────────────────────────────────────┐
│ TODAY — "best next session" recommendation │
└─────────────────────────────────────────────────────────┘

┌─────────────────────────────────────────────────────────┐
│ MEMORY MODEL (EchoType FSRS — keep, keep separate) │
│ ts-fsrs scheduling. Memory strength ≠ ability. │
└─────────────────────────────────────────────────────────┘
```

## Capability model (kernel-owned)

- Capability = observable ability under stated conditions
(`reception.*` / `production.*` / `interaction.*`), never a lesson
position. Prerequisites = hard dependency edges only; pedagogy
ordering lives in mission `recommendedAfterMissions`.
- Claims are milestone states: `NOT_SEEN → EXPOSED → SUPPORTED →
INDEPENDENT → RETAINED → TRANSFERRED`, each with its own evidence
bar. `assessment` context ≠ `transfer` context.

## Task & mission contracts (kernel-owned)

- Task purposes: diagnostic, input, notice, retrieval, production,
interaction, remediation, delayed_retrieval, transfer, assessment,
support. Purpose gates the event types it may mint
(`EVENT_TYPES_FOR_PURPOSE`) — a retrieval task cannot emit
`transfer_attempt`.
- Missions combine capabilities in a real-world scenario with a staged
lifecycle (diagnostic → input → supported attempt → feedback →
retry → delayed retrieval → transfer → assessment).
- Contracts are executable validators — a task violating its own
purpose class fails before learners touch it.

## Evidence flow

```
UI action → outcome observed (score/transcript/self-report)
→ evaluator (authority: deterministic|human|asr|ai_llm|self_report)
→ bindAttempt(task, capability, raw) ← sole mint, throws on
caller-supplied semantics (FORGED_FIELDS)
→ EvidenceEvent (validated, support flags, binding provenance)
→ append to event log (Dexie events table; Supabase sync later)
→ projectLearnerState / learnerModel / planNext
```

Rules that make this trustworthy:

- `asr`/`ai_llm`/`self_report` authority can never award independent
ability — STT transcript match cannot mint pronunciation evidence.
- Support flags (hint/translation/transcript/modelAnswer/repeat) are
honest provenance: recorded always, credited by conditions —
`answerBearing` support demotes an attempt to SUPPORTED.
- `unionSupport` accumulates per attemptId — a retry cannot launder a
hinted attempt into "unaided".
- Events are immutable, deduped by id, replayed in canonical order —
learner state is always a projection, never stored.

## Planner / Next For You

- `planNext`/`selectNextTask` produce the next task **with a decision
record** (`decisionAuditRecord`) — "12 min — best next session"
carries a reason the UI can show.
- Support-demand routing: a failed attempt whose evaluation reports
`missingFunctions` can mint a `support` probe — remediation context,
never claim-bearing.
- Correction episodes track repair lifecycle (miss → repair → retest
→ verified/relapsed) with burned-surface semantics.

## Speaking evaluation boundary

- STT (Web Speech / server fallback) records response + history only.
- `evaluation.authority = 'asr'` — response recorded, never
pronunciation/transfer credit.
- Candidate acoustic evaluator: Halleck45/OpenPronounce (MIT) —
expected/heard phones, phoneme diffs, word errors, confidence,
prosody. **Audit before integrate**: model size, latency, hosting,
correctness vs a labeled sample. Not in milestone 1.

## Content / import flow

Keep EchoType's pipeline (YouTube/PDF/URL/audio/text) → `contents`
table. Kernel overlay: imported content becomes carrier material;
curriculum authoring binds tasks to capabilities via
`curriculum-checks` (namespaced ids, `targetCapabilities` ≤3 with
coverage debts, `carrierCapabilities`, `supportCapabilities`,
`contextSignature`, canonical `pf.<cap>.<cue>.<setting>.<register>.
<channel>.<hash>.vN` prompt families).

## Persistence strategy

- **Events (kernel)**: Dexie table `evidenceEvents` (schema v21:
`id, learnerId, taskId, capabilityId, occurredAt,
[learnerId+occurredAt]`) — append-only by convention + content-
fingerprint dedupe (identical redelivery dedupes; same id +
different content throws). Supabase mirror: append-only table +
insert-only policy. `occurredAt` (client-observed) vs
`recordedAt` (server-arrival) kept distinct.
- **Everything else**: existing EchoType tables unchanged.
- FSRS data stays on `records.fsrsCard` — scheduler state, not ability.

## Migration strategy

1. DONE — vendored kernel byte-identically into `src/vnext/` (26
files) + the one pure external dep `src/core/mission-checks.js`
(imported by `evaluators.js`). `mission-runner.js` IS vendored —
`selectNextTask`'s shipped `reference` mode is `nextMissionTask`.
Excluded: `ui-session.js`, `ui/` (FlashDay-UI coupled —
`../../ui/speech.js`). Vendored code is biome-excluded and
typecheck-unchecked (`include` covers `.ts`/`.tsx` only) so a
future kernel sync is a byte-diff, never a merge conflict.
Sanity: 18/18 modules import cleanly in Node; `allowJs` resolves
the `.js` imports for TypeScript consumers.
2. DONE — bridge shipped at `src/lib/evidence-bridge/`:
- `registry.ts` — `createRegistry` runs the kernel authoring gate
(`checkCurriculum`) so an unshippable contract set cannot mint
evidence; `fixtureRegistry()` wires the vendored curriculum.
- `bridge.ts` — `submitAttempt`/`submitObservation` accept only
observed reality (taskId + response/support/timing); forged
semantic keys throw; deterministic `contractId` tasks re-score
the response via `evaluateAttempt` (caller outcome ignored).
- `store.ts` — `createDexieEventStore` mirrors the kernel store
surface over `db.evidenceEvents`.
- `projectState` / `nextAction` — replay-derived learner state
and Next-For-You decisions with reasons.
- 27 contract pins in `evidence-bridge.test.ts` (forgery,
correct!=capability, completion!=mastery, supported!=independent,
STT!=pronunciation, practiced!=transfer, assessment attemptId,
FSRS-inertness, learner isolation, dedupe/conflict, persist->
replay determinism, retention lag, auditable decisions).
3. Today surface: render `selectNextTask` decision + reason.
4. Vertical slice: `meeting-someone` mission through real UI —
`mission.meet_new_person` fixtures already carry 18 tasks.
5. Rebrand only after slice verified — name at the shell surface,
MIT attribution preserved in `LICENSE` + NOTICE.

## Non-goals (per mission)

- No A1→C2 curriculum; initial guided path A0/A1 → early A2.
- No EchoType rewrite; no mass renaming.
- No LLM-authored proficiency; no gamification-as-evidence.
- No STT-as-pronunciation; no FSRS-as-ability.
Loading