Skip to content

feat(article): falsification the schema judges, not a heuristic - #52

Closed
Willi363363 wants to merge 1 commit into
feat/rewrite-phase-3-mediawikifrom
feat/rewrite-phase-3-falsify
Closed

Willi363363 wants to merge 1 commit into
feat/rewrite-phase-3-mediawikifrom
feat/rewrite-phase-3-falsify

Conversation

@Willi363363

Copy link
Copy Markdown
Owner

Step 3.4 of plans/rewrite/phase-03-article.md.

Stacked on #51. Base is feat/rewrite-phase-3-mediawiki, so the diff here is
step 3.4 alone.

130 lines of heuristics collapse into one line

They were business logic by accident, and every one existed because the answer
was a string:

heuristic replaced by
strip ```json fences the schema
fall back from first [ to last ] the schema
unwrap an envelope object the schema
all-or-nothing policy on indices the schema
positional fallback when indices did not match dropping the item
a second request for whatever the first lost one failure path

Six malformed shapes are asserted rejected — prose around the JSON, a Markdown
fence, a bare array, not-JSON-at-all, an empty answer, an envelope with the wrong
key — each of which had its own heuristic.

The sheet names a deprecated API

It says generateObject. That function is deprecated in the installed
version: the SDK moved to generateText with Output.object(). Same guarantee,
different call.

I checked the installed package rather than trusting what I remembered of the
SDK, and the sheet now says so — the next reader would otherwise reach for a
deprecated function on its word.

The 1000-character truncation is fixed

Today the model receives text[:1000] while the player is served the paragraph
in full. So the model rewrites an ending it never saw, and the second half of
the paragraph contradicts the first half the player is being graded on.

Verified biting: reintroducing .slice(0, 1000) fails
sends a 2000-character paragraph whole.

The positional fallback is gone, not ported

A model that quotes an index it was never given is not partially right — it is
describing a paragraph nobody will be graded on, and the current code turns that
into a wrong grade. Those falsifications are dropped; an answer with nothing
usable is a failure.

Verified biting: reintroducing the fallback fails two tests.

Also collapsed: two falsifications on one paragraph. C3.3 forbids it and the
game_position unique constraint from #45 refuses to store it, so keeping one is
what stops a model quirk from becoming a database error.

The prompt is verbatim

In French, unchanged, from misinformation.py — not the dead one in
core/prompts.py. This phase's pitfall is explicit: Output.object may already
move the model's output, and changing the wording in the same step would make it
impossible to know which did it. A test asserts the wording is still there.

One constant instead of two

The falsifiability floor exists twice today: MIN_FALSIFIABLE_CHARS in settings
and MIN_PARAGRAPH_LENGTH = 100 hard-coded in misinformation.py. That is the
duplication D8 names and this phase's last pitfall asks to close — now one
exported constant, with a test pinning it.

Checks

CI=true pnpm test   article 82 · db 61 · domain 233 · protocol 220 · config 19 · env 7
pnpm typecheck      6 packages
pnpm lint           6 packages
pnpm format:check
bash scripts/checks.sh staged

Two new dependencies: ai (the SDK named by the decision table) and zod
(already used across the monorepo). The model is a parameter, not a provider
chosen here — so no @ai-sdk/* package yet, the tests use MockLanguageModelV4,
and phase 4 decides where the model comes from.

Not run: the Python backend tests — no pytest here. This touches no Python.

Next

3.5 (end-to-end parity on fixtures with a mocked model), 3.6 (Redis cache),
3.7 (counters).

🤖 Generated with Claude Code

https://claude.ai/code/session_01533WTTbgmfCaDP5tmK8Gkv

Step 3.4 of phase 3. About 130 lines of parsing heuristics disappear, and
they were business logic by accident: stripping Markdown fences, falling back
from the first `[` to the last `]` when the JSON did not parse, unwrapping an
envelope object, an all-or-nothing policy on indices, a positional fallback
when the model renumbered them, and a second request for whatever the first
one lost.

Every one of those existed because the answer was a string. Ask for an object
and they collapse into one line: the schema validates or it does not. Six
malformed shapes are asserted rejected, each of which had its own heuristic.

The sheet names `generateObject`, and that API is **deprecated** in the
version installed — the SDK moved to `generateText` with `Output.object()`.
Same guarantee, different call. The sheet says so now, because the next
reader would otherwise reach for a deprecated function on its word. I checked
the installed package rather than trusting what I remembered of the SDK.

The 1000-character truncation is fixed. Today the model receives
`text[:1000]` while the player is served the paragraph in full: the model
rewrites an ending it never saw, and the second half of the paragraph
contradicts the first. Verified biting — reintroducing the slice fails the
test.

The positional fallback is gone rather than ported. A model that quotes an
index it was never given is not partially right: it is describing a paragraph
nobody will be graded on. Those are dropped, and an answer with nothing
usable is a failure. Verified biting.

The prompt is carried over verbatim, in French, because this phase's pitfall
says not to mix a stack change with a behaviour change: `Output.object` may
already move the model's output, and changing the wording in the same step
would make it impossible to know which did.

The falsifiability floor is one constant now. It exists twice today —
`MIN_FALSIFIABLE_CHARS` in settings and `MIN_PARAGRAPH_LENGTH = 100`
hard-coded in `misinformation.py` — which is the duplication D8 names and this
phase's last pitfall asks to close.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01533WTTbgmfCaDP5tmK8Gkv
@qodo-code-review

Copy link
Copy Markdown

PR Summary by Qodo

Use AI SDK structured output + Zod to validate falsification responses

✨ Enhancement 🧪 Tests 📝 Documentation ⚙️ Configuration changes 🕐 40+ Minutes

Grey Divider

AI Description

• Switch falsification parsing to schema-validated structured output (no string heuristics).
• Stop truncating paragraphs sent to the model; keep indices strict and dedupe per paragraph.
• Add deterministic unit tests and update rewrite plan to reflect generateText/Output.object().
Diagram

graph TD
A["Article pipeline"] --> B["falsify.ts"] --> C{{"AI SDK generateText"}} --> D{{"LLM provider"}} --> E["Structured object" ] --> F["Index filter & dedupe"] --> G["Result (ok/failed)"]
H[["falsify.test.ts (MockLanguageModel)" ]] --> B
B --> I["Zod schema"] --> C
subgraph Legend
  direction LR
  _mod["Module"] ~~~ _ext{{"External/SDK"}} ~~~ _test[["Test"]]
end
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Keep string parsing + heuristics
  • ➕ No new runtime dependencies (ai/zod).
  • ➕ Potentially more tolerant of sloppy model outputs.
  • ➖ Heuristics become implicit business logic and are hard to reason about.
  • ➖ Increases risk of accepting malformed/shifted outputs (grading wrong paragraph).
  • ➖ Multiple partial-failure paths; harder to test exhaustively.
2. Parse JSON first, then validate with Zod (without Output.object)
  • ➕ Works even without SDK structured-output support.
  • ➕ Clear separation: parsing vs validation.
  • ➖ Still needs recovery rules for markdown fences/prose envelopes/truncation.
  • ➖ Reintroduces the core problem: ambiguous strings require heuristics.
3. Use provider function/tool calling instead of Output.object
  • ➕ Often yields reliably typed payloads with some providers.
  • ➕ Can encode stricter constraints at the provider layer.
  • ➖ More provider-specific surface area and configuration complexity.
  • ➖ Harder to keep consistent across model backends than SDK-level structured output.

Recommendation: The PR’s approach (AI SDK generateText + Output.object() + Zod) is the best fit because it collapses many brittle, business-logic-by-accident heuristics into a single validation gate with one failure path. The added tests meaningfully lock in the intended semantics (no truncation, strict indices, dedupe), which is the main risk area when changing LLM integration behavior.

Files changed (7) +533 / -7

Enhancement (2) +196 / -1
falsify.tsImplement falsify() via generateText + Output.object() with strict schema +183/-0

Implement falsify() via generateText + Output.object() with strict schema

• Adds a new falsification module that requests structured output validated by a Zod schema, replacing prior string parsing heuristics with a single validation/failure path. Ensures full paragraph text is sent (no 1000-char truncation), drops falsifications with unoffered indices, enforces one falsification per paragraph, sorts results, and returns provider token usage as nullable values.

packages/article/src/falsify.ts

index.tsExport falsification API surface from article package +13/-1

Export falsification API surface from article package

• Re-exports falsification functions, constants, and types so downstream packages can call 'falsify()' and compute 'falsifiableCandidates()' with shared constants.

packages/article/src/index.ts

Tests (2) +241 / -1
falsify.test.tsAdd deterministic falsify() test suite with mocked model +240/-0

Add deterministic falsify() test suite with mocked model

• Introduces unit tests validating candidate selection, no paragraph truncation, schema-only acceptance/rejection, dropping invented indices, per-paragraph dedupe, prompt wording stability, and usage accounting. Uses 'MockLanguageModelV4' so failures are deterministic and tied to integration logic rather than model behavior.

packages/article/src/falsify.test.ts

workspace-graph.test.tsUpdate workspace dependency expectations for article package +1/-1

Update workspace dependency expectations for article package

• Adjusts the workspace dependency graph test to reflect the article package’s new dependencies on 'ai' and 'zod'.

packages/config/src/workspace-graph.test.ts

Documentation (1) +11 / -4
phase-03-article.mdDocument structured-output approach and deprecated generateObject note +11/-4

Document structured-output approach and deprecated generateObject note

• Updates step 3.4 to describe structured output using 'generateText' with 'Output.object()' (not deprecated 'generateObject') and records that the falsifiability floor is now a single constant ('MIN_FALSIFIABLE_CHARS') in the target package.

plans/rewrite/phase-03-article.md

Other (2) +85 / -1
package.jsonAdd AI SDK + Zod runtime dependencies for structured output +3/-1

Add AI SDK + Zod runtime dependencies for structured output

• Adds 'ai' and 'zod' dependencies to support structured output generation and schema validation inside the article package.

packages/article/package.json

pnpm-lock.yamlLock new AI SDK transitive dependencies and Zod version +82/-0

Lock new AI SDK transitive dependencies and Zod version

• Updates the lockfile to include 'ai@7.0.77', 'zod@4.4.3', and required transitive packages for structured output support.

pnpm-lock.yaml

@qodo-code-review

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (1) 📘 Rule violations (0) 📜 Skill insights (0)

Grey Divider


Action required

1. Prompt schema contract conflicts 🐞 Bug ≡ Correctness
Description
systemPrompt instructs the model to output a bare JSON array with snake_case keys like
paragraph_index, but Output.object validates a camelCase { falsifications: [...] } envelope
with fields like paragraphIndex. As a result, a model that follows the prompt fails Zod validation
and falsify() returns unexpected_response, while tests mask the mismatch by mocking the
schema-shaped response the prompt never requests.
Code

packages/article/src/falsify.ts[R80-84]

+Tu dois retourner un tableau JSON d'objets (un par paragraphe dans le même ordre) avec exactement ces clés :
+- "paragraph_index": l'indice d'origine fourni dans la requête (un entier).
+- "swapped_text": Le paragraphe complet modifié.
+- "explanation": Une explication très courte (1 phrase) sur LA VÉRITÉ.
+- "hint": Un indice très court pour aider le joueur (ex: "Vérifiez cette date d'élection").
Evidence
The system prompt explicitly requests a raw JSON array using snake_case fields (e.g.,
paragraph_index), while the structured-output configuration uses Output.object({ schema }) with
a Zod schema that requires a top-level object envelope { falsifications: [...] } and camelCase
fields like paragraphIndex. Because Output.object validates the generated data against the
provided schema (and does not auto-wrap arrays or rename fields), any response that adheres to the
prompt will fail parsing/validation; the implementation catches that failure and converts it to
unexpected_response. The tests further obscure the issue by mocking/expecting the envelope and
camelCase form rather than the exact response contract described in the prompt.

packages/article/src/falsify.ts[41-48]
packages/article/src/falsify.ts[80-86]
packages/article/src/falsify.ts[142-157]
packages/article/src/falsify.test.ts[42-56]
backend/src/core/misinformation.py[136-157]
🌐 AI SDK 7 documents that Output.object generates and validates data against the supplied schema, and generateText rejects when the response cannot be parsed or validated.

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The prompt’s required JSON shape and field names contradict the Zod schema used for structured output (`Output.object`). Make the prompt and schema specify the same response contract (shape + casing), and only map/rename fields after successful validation if the public TypeScript API needs a different representation; also adjust tests so they verify the full requested response shape rather than masking the mismatch via mocks.

## Issue Context
The carried-over Python-style prompt asks for a bare array with snake_case keys such as `paragraph_index`, `swapped_text`, etc., but the new structured-output schema expects a top-level `{ falsifications: [...] }` envelope with camelCase fields (e.g., `paragraphIndex`). `Output.object` validates model output against the Zod schema and does not rename keys or wrap arrays; schema-validation failures are caught and returned as `unexpected_response`. Current tests hide the discrepancy by mocking/expecting the envelope + camelCase shape that the prompt never requests.

## Fix Focus Areas
- packages/article/src/falsify.ts[41-48]
- packages/article/src/falsify.ts[72-107]
- packages/article/src/falsify.ts[142-145]
- packages/article/src/falsify.test.ts[128-160]
- packages/article/src/falsify.test.ts[207-219]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context sources
✅ Web pages:
  +2 more
Review mode: 🧠 Deep: This introduces substantial new structured-output/runtime logic, prompt and dependency changes, and integration-facing behavior across multiple independent paths, making multiple subtle defects plausible.

Grey Divider

Tip of the day
💡 Did you know, you can switch off images and animations for a plain-text comment

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment on lines +80 to +84
Tu dois retourner un tableau JSON d'objets (un par paragraphe dans le même ordre) avec exactement ces clés :
- "paragraph_index": l'indice d'origine fourni dans la requête (un entier).
- "swapped_text": Le paragraphe complet modifié.
- "explanation": Une explication très courte (1 phrase) sur LA VÉRITÉ.
- "hint": Un indice très court pour aider le joueur (ex: "Vérifiez cette date d'élection").

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Action required

1. Prompt schema contract conflicts 🐞 Bug ≡ Correctness

systemPrompt instructs the model to output a bare JSON array with snake_case keys like
paragraph_index, but Output.object validates a camelCase { falsifications: [...] } envelope
with fields like paragraphIndex. As a result, a model that follows the prompt fails Zod validation
and falsify() returns unexpected_response, while tests mask the mismatch by mocking the
schema-shaped response the prompt never requests.
Agent Prompt
## Issue description
The prompt’s required JSON shape and field names contradict the Zod schema used for structured output (`Output.object`). Make the prompt and schema specify the same response contract (shape + casing), and only map/rename fields after successful validation if the public TypeScript API needs a different representation; also adjust tests so they verify the full requested response shape rather than masking the mismatch via mocks.

## Issue Context
The carried-over Python-style prompt asks for a bare array with snake_case keys such as `paragraph_index`, `swapped_text`, etc., but the new structured-output schema expects a top-level `{ falsifications: [...] }` envelope with camelCase fields (e.g., `paragraphIndex`). `Output.object` validates model output against the Zod schema and does not rename keys or wrap arrays; schema-validation failures are caught and returned as `unexpected_response`. Current tests hide the discrepancy by mocking/expecting the envelope + camelCase shape that the prompt never requests.

## Fix Focus Areas
- packages/article/src/falsify.ts[41-48]
- packages/article/src/falsify.ts[72-107]
- packages/article/src/falsify.ts[142-145]
- packages/article/src/falsify.test.ts[128-160]
- packages/article/src/falsify.test.ts[207-219]

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools

@Willi363363

Copy link
Copy Markdown
Owner Author

Closed by the stack collapse, not abandoned.

Every commit of this pull request is in willi363/refonte via #114, which
merged the whole rewrite in one go: the stack was strictly linear, so the top
branch was an ancestor-of-nothing and a descendant-of-everything, and the merge
was a fast-forward with no conflict possible. The history keeps one commit per
step, which is what squash-per-step was there to produce.

Three commit messages were reworded on the way — a wip: type, a 74-character
subject and a (web,realtime) scope, none of which scripts/checks.sh allows.
The trees are byte-identical.

This description and its review thread stay readable here.

@Willi363363
Willi363363 deleted the feat/rewrite-phase-3-falsify branch August 29, 2026 16:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant