Repository navigation
fix(server): preserve stored effort for agent-started turns - #134
Conversation
Server-origin turn commands omitted modelSelection, so the provider reactor sent turns without it and Codex fell back to medium reasoning effort. The shared fork turn sender (threads toolkit, Wight, usage-limit resume, scheduler), the Spectrum turn builder (Colors, moderator, caller report) and the Prism stale-child notice now carry the target thread's stored modelSelection. Closes #133 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
lukemaj
left a comment
There was a problem hiding this comment.
Independent exact-head review: CLEAN at a4fee31 (base 86d3269).
Findings: none requiring a change in #133 scope. The shared sender addition covers spawn, ordinary message, child report, Wight, usage-limit resume, and scheduler turns. Spectrum has its own builder; its stored selection is used for each participant turn, moderator synthesis, and caller report, and outbox replay spreads the command. The stale-child notice uses its parent thread selection. The test assertions cover Spectrum participants and moderator, plus caller report, and the spawned Codex child reaches the ProviderService sendTurn boundary through the real engine/reactor.
Inspected the complete six-file diff, reachable fork dispatch sites, decider turn-start persistence, reactor selection logic, and compaction replay. Compaction replay spreads the persisted turn-start payload, which includes modelSelection. The async-answer path is upstream-owned and user-initiated; per the stated scope it is not a blocker. Claude source evidence supports the handoff: initial SDK query effort comes from startSession; sendTurn does not set live SDK effort, while selected effort can affect the prompt-injected ultrathink prefix.
Proof inspected (not rerun by this reviewer): #133 handoff records focused failing-before results (3 failed/26 passed with source lines reverted), passing-after results (3 files/29 passed), adapter receipt plus reactor drain without sleeps, targeted module tests, changed-file lint/format, server typecheck, and fork-check. No live verification is claimed.
Agent work on this PREstimated cost unknown · 0 responses · 64 sessions · 5.0 h wall time
Flags: 6 human corrections · 115 large tool outputs · 1 permission request · 62 repeated commands · 15 repeated failures · 36 repeated reads · 11 repeated skill loads · 63 sessions with usage bound to no task · 1 session without usage records Details: snapshot, prices, coverage, counters
Token counters by model (native counter semantics; never added across semantics): Selected rates (USD per million tokens). These rates value the report at the selected schedule date; they do not establish historical prices or subscription spending.
Other output and reasoning are priced without double counting inclusive native output. Missing rates remain unknown. Local measurement from native records; usage totals are not billing. Updated in place by |
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
What: Agent-started turns now preserve each thread's stored model and reasoning effort.
Why: Fork turn builders omitted
modelSelection, causing Codex to default tomediumregardless of the configured Color.So what: Independent review passed on the exact head; the user merges this single PR once GitHub CI is green. No installation is part of this task.
Closes #133. References #113 and #119 and restores the configured Colors required by the Promachos Objective.
Three fork-owned field additions fix the shared sender (spawn, messages, child reports, Wight, usage-limit resume and scheduler), Spectrum's participant/moderator/caller-report builder, and Prism's stale-parent notifier. No refactor machinery, upstream edits, UI changes or provider redesign. No versioning by design.
Audit and Claude behavior
The Issue proof and durable handoff records every dispatch path and source reference. Promachos startup already supplies its Prism selection. Compaction and Spectrum outbox replay spread the persisted request and retain the original selection.
Async-answer synthesis in
orchestration/decider.ts:1917is upstream-owned, identical at upstream merge-basede251fcand inspected upstream5cc99e1. It is a user-initiated path and remains unchanged, outside this fork-origin scope. Codex message-mode answers remain an existing upstream limitation, not an acceptance blocker for this fork-origin fix.Claude has no equivalent fallback to medium: startup supplies the stored selection and creates the SDK query with its configured effort (
ClaudeAdapter.ts:4828-4843,4914-4918). Without per-turn selection,sendTurnpreserves that query's startup effort and leavesstartInput/currentEffortunchanged. Supplying selection updates metadata but calls no SDK effort setter, so this PR does not claim to change a running query's effort. It restores the existingUltrathink:prompt-prefix behavior for agent-started turns. Changing stored model/effort during an existing Claude session is a separate non-goal.Proof on a4fee31
Run from
apps/serverunless stated:pnpm exec vp test run src/mcp/toolkits/threads/sendThreadTurn.test.ts src/mcp/toolkits/threads/spectrum.test.ts src/prism/staleTurnMonitor.test.ts: before the three production additions, 3 failed, 26 passed, all three failures show omitted selection; after, 29 passed.pnpm exec vp test run src/mcp/toolkits/threads/ src/scheduler/ src/prism/staleTurnMonitor.test.ts: 170 passed, 20 files.pnpm exec vp lint --report-unused-disable-directives <six changed files>andpnpm exec vp fmt --check <six changed files>: pass.pnpm exec tsc --noEmit: pass.scripts/fork-check.sh: pass; all six changed paths are fork-owned.The spawned-child regression invokes
spawn_threadthrough the real engine, projections and ProviderCommandReactor to a faked provider boundary; a Deferred send receipt and reactor drain gate its assertion of storedreasoningEffort: xhigh. Spectrum settlement receipts prove participant, moderator and caller-report selections. The stale-notice test asserts the parent's selection. No sleeps, repo-wide checks, browser, dev server, live-state writes or installed-build verification.Independent review
Exact candidate:
a4fee3127d2d29227ecf694237894282390363d9.Reviewer:
sub.sub.sub.46b03420-8643-4703-8756-8602f14b39fb.dispatcher-744df2c8289e.dispatcher-a529762782a5.reviewer-00eefad4d0ca, Prism-routed Luna via Codex.Verdict: CLEAN, no acceptance-blocking findings. Exact-head review; newest
review/independentstatus is success on this SHA, creatorlukemaj. The reviewer inspected the complete diff and proof, without rerunning the targeted checks.GitHub CI at handoff: Check and all three Test Server shards pass on this exact head; the general Test job remains running. CI run. No CI failure is reported.
Cost evidence
Agent Observer publication is posted. Attribution is incomplete: it reports no attributed responses, 63 sessions with usage bound to no task, and one without usage records. Complete cost is unknown, not zero; inherited/shared-session totals are not treated as this PR's cost.
Reversible assumptions and limits
Requirements and who asked: User and #133 require stored effort across every fork-origin dispatch path, Spectrum participants/moderator, stale notices, explicit upstream audit decisions, precise Claude evidence, receipt-based red/green proof and one independently reviewed PR.
Deleted: Per-caller patches beyond the three builders, refactor machinery, upstream/provider edits, UI, versioning, repo-wide checks, merge and installation.
Bottleneck: Fork turn builders dropped
modelSelection, so supported in-session adapters received no turn selection and Codex defaulted to medium.Checked myself: Complete exact-head diff, pushed branch identity, three fork-owned production additions, receipt/drain assertions, source audit and Issue proof; independent review and newest
review/independent=successverified on this exact SHA.Implementation: Claude Opus 5.5 via Claude Code in T3 Code. Coordination: GPT-6.1-Sol via Codex in T3 Code. Independent review: Prism-routed Luna via Codex.