Skip to content

fix(server): keep runs alive after malformed provider snapshots - #14211

Closed
saphid wants to merge 1 commit into
pingdotgg:t3code/codex-turn-mappingfrom
saphid:agent/v2-ingest-nonfatal-upstream
Closed

saphid wants to merge 1 commit into
pingdotgg:t3code/codex-turn-mappingfrom
saphid:agent/v2-ingest-nonfatal-upstream

Conversation

@saphid

@saphid saphid commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

A schema-invalid provider snapshot currently ends root-run ingestion even when the provider can continue. This change logs and skips schema-rejected message, turn-item, node, subagent, plan, provider-thread, and provider-turn snapshots. Later valid snapshots can restore those entities while ingestion remains active.

Terminal, thread-creation, and runtime-request errors retain the existing failure path, as do storage failures, other normalization errors, defects, and interruption. Rejected snapshots do not create background ownership or extend ingestion past root completion.

Recovery is best effort: skipped snapshots have no durable rejection record and may leave missing output, stale state, or pending child prompts without a valid replacement. A rejected settled snapshot for an already-tracked child can retain its subscription until the session is released. Warnings are logged per rejection. There is no confirmed production schema-rejection incident; the motivating SQLite failure is addressed separately by #14365.

Verified with 171 tests across four focused files, server typecheck, and lint/format checks on the three changed files. Regression tests reproduce malformed lone-terminal, child-creation, and runtime-request failures before the repair and pass afterward. Independent review reran the 70 ingestor/execution tests and found no blocking code findings. This backend-only change has no visual behavior to demonstrate.

Current CI limitation: six jobs in run 36727099830 failed during Vite+ installation after four attempts, before their checks/tests ran. Rerunning failed jobs was rejected because this account lacks repository admin rights. The PR remains a draft pending a successful current-head CI run.

Readiness repairs: GPT-6 Astra in the Codex harness through T3 Code. Independent review: Anthropic Claude Fable 5.1, high reasoning, Claude Code harness through T3 Code.

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:M 30-99 changed lines (additions + deletions). labels Sep 29, 2026
@saphid
saphid force-pushed the agent/v2-ingest-nonfatal-upstream branch from 6761b7f to 371bc63 Compare September 30, 2026 14:10
@saphid saphid changed the title fix(server): one failed provider event no longer fails a running run fix(server): keep runs alive after malformed provider snapshots Sep 30, 2026

Copy link
Copy Markdown
Member

Note

This comment is posted by Julius' dot

This changes schema-rejected snapshots from a failed run into best-effort continuation, which can leave missing output, stale state or pending child prompts. The reported SQLite incident is handled separately in #14365. Could a maintainer confirm that this skip-and-continue policy is the intended direction, and its acceptable limits? Please link that decision under prior approval.

@juliusmarminge
juliusmarminge deleted the branch pingdotgg:t3code/codex-turn-mapping October 2, 2026 19:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M 30-99 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants