Skip to content

fix(server): preserve stored effort for agent-started turns - #134

Merged
lukemaj merged 1 commit into
mainfrom
fix/turn-effort
Oct 2, 2026
Merged

lukemaj merged 1 commit into
mainfrom
fix/turn-effort

Conversation

@lukemaj

@lukemaj lukemaj commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

What: Agent-started turns now preserve each thread's stored model and reasoning effort.
Why: Fork turn builders omitted modelSelection, causing Codex to default to medium regardless of the configured Color.
So what: Independent review passed on the exact head; the user merges this single PR once GitHub CI is green. No installation is part of this task.

Closes #133. References #113 and #119 and restores the configured Colors required by the Promachos Objective.

Three fork-owned field additions fix the shared sender (spawn, messages, child reports, Wight, usage-limit resume and scheduler), Spectrum's participant/moderator/caller-report builder, and Prism's stale-parent notifier. No refactor machinery, upstream edits, UI changes or provider redesign. No versioning by design.

Audit and Claude behavior

The Issue proof and durable handoff records every dispatch path and source reference. Promachos startup already supplies its Prism selection. Compaction and Spectrum outbox replay spread the persisted request and retain the original selection.

Async-answer synthesis in orchestration/decider.ts:1917 is upstream-owned, identical at upstream merge-base de251fc and inspected upstream 5cc99e1. It is a user-initiated path and remains unchanged, outside this fork-origin scope. Codex message-mode answers remain an existing upstream limitation, not an acceptance blocker for this fork-origin fix.

Claude has no equivalent fallback to medium: startup supplies the stored selection and creates the SDK query with its configured effort (ClaudeAdapter.ts:4828-4843,4914-4918). Without per-turn selection, sendTurn preserves that query's startup effort and leaves startInput/currentEffort unchanged. Supplying selection updates metadata but calls no SDK effort setter, so this PR does not claim to change a running query's effort. It restores the existing Ultrathink: prompt-prefix behavior for agent-started turns. Changing stored model/effort during an existing Claude session is a separate non-goal.

Proof on a4fee31

Run from apps/server unless stated:

  • pnpm exec vp test run src/mcp/toolkits/threads/sendThreadTurn.test.ts src/mcp/toolkits/threads/spectrum.test.ts src/prism/staleTurnMonitor.test.ts: before the three production additions, 3 failed, 26 passed, all three failures show omitted selection; after, 29 passed.
  • pnpm exec vp test run src/mcp/toolkits/threads/ src/scheduler/ src/prism/staleTurnMonitor.test.ts: 170 passed, 20 files.
  • Root pnpm exec vp lint --report-unused-disable-directives <six changed files> and pnpm exec vp fmt --check <six changed files>: pass.
  • Server pnpm exec tsc --noEmit: pass.
  • Root scripts/fork-check.sh: pass; all six changed paths are fork-owned.

The spawned-child regression invokes spawn_thread through the real engine, projections and ProviderCommandReactor to a faked provider boundary; a Deferred send receipt and reactor drain gate its assertion of stored reasoningEffort: xhigh. Spectrum settlement receipts prove participant, moderator and caller-report selections. The stale-notice test asserts the parent's selection. No sleeps, repo-wide checks, browser, dev server, live-state writes or installed-build verification.

Independent review

Exact candidate: a4fee3127d2d29227ecf694237894282390363d9.
Reviewer: sub.sub.sub.46b03420-8643-4703-8756-8602f14b39fb.dispatcher-744df2c8289e.dispatcher-a529762782a5.reviewer-00eefad4d0ca, Prism-routed Luna via Codex.
Verdict: CLEAN, no acceptance-blocking findings. Exact-head review; newest review/independent status is success on this SHA, creator lukemaj. The reviewer inspected the complete diff and proof, without rerunning the targeted checks.

GitHub CI at handoff: Check and all three Test Server shards pass on this exact head; the general Test job remains running. CI run. No CI failure is reported.

Cost evidence

Agent Observer publication is posted. Attribution is incomplete: it reports no attributed responses, 63 sessions with usage bound to no task, and one without usage records. Complete cost is unknown, not zero; inherited/shared-session totals are not treated as this PR's cost.

Reversible assumptions and limits

  • Use the target thread's stored selection rather than a separate Spectrum participant copy.
  • Other adapters receive the same stored selection they already receive for client-origin turns; they were not separately exercised.
  • Existing user guidance stays accurate; no documentation or version changes are needed.
  • Async-answer synthesis remains upstream scope; installed-build native-effort confirmation and rerunning Compare Spectrum modes with judged scenarios #119 are outside this task.
  • Website/SEO impact: none, because this is server turn-command preservation with no website or URL change.

Requirements and who asked: User and #133 require stored effort across every fork-origin dispatch path, Spectrum participants/moderator, stale notices, explicit upstream audit decisions, precise Claude evidence, receipt-based red/green proof and one independently reviewed PR.
Deleted: Per-caller patches beyond the three builders, refactor machinery, upstream/provider edits, UI, versioning, repo-wide checks, merge and installation.
Bottleneck: Fork turn builders dropped modelSelection, so supported in-session adapters received no turn selection and Codex defaulted to medium.
Checked myself: Complete exact-head diff, pushed branch identity, three fork-owned production additions, receipt/drain assertions, source audit and Issue proof; independent review and newest review/independent=success verified on this exact SHA.

Implementation: Claude Opus 5.5 via Claude Code in T3 Code. Coordination: GPT-6.1-Sol via Codex in T3 Code. Independent review: Prism-routed Luna via Codex.

Server-origin turn commands omitted modelSelection, so the provider
reactor sent turns without it and Codex fell back to medium reasoning
effort. The shared fork turn sender (threads toolkit, Wight, usage-limit
resume, scheduler), the Spectrum turn builder (Colors, moderator, caller
report) and the Prism stale-child notice now carry the target thread's
stored modelSelection.

Closes #133

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XS labels Oct 2, 2026

@lukemaj lukemaj left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Independent exact-head review: CLEAN at a4fee31 (base 86d3269).

Findings: none requiring a change in #133 scope. The shared sender addition covers spawn, ordinary message, child report, Wight, usage-limit resume, and scheduler turns. Spectrum has its own builder; its stored selection is used for each participant turn, moderator synthesis, and caller report, and outbox replay spreads the command. The stale-child notice uses its parent thread selection. The test assertions cover Spectrum participants and moderator, plus caller report, and the spawned Codex child reaches the ProviderService sendTurn boundary through the real engine/reactor.

Inspected the complete six-file diff, reachable fork dispatch sites, decider turn-start persistence, reactor selection logic, and compaction replay. Compaction replay spreads the persisted turn-start payload, which includes modelSelection. The async-answer path is upstream-owned and user-initiated; per the stated scope it is not a blocker. Claude source evidence supports the handoff: initial SDK query effort comes from startSession; sendTurn does not set live SDK effort, while selected effort can affect the prompt-injected ultrathink prefix.

Proof inspected (not rerun by this reviewer): #133 handoff records focused failing-before results (3 failed/26 passed with source lines reverted), passing-after results (3 files/29 passed), adapter receipt plus reactor drain without sleeps, targeted module tests, changed-file lint/format, server typecheck, and fork-check. No live verification is claimed.

@lukemaj

lukemaj commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

Agent work on this PR

Estimated cost unknown · 0 responses · 64 sessions · 5.0 h wall time

Model Responses Tokens Estimated cost

Flags: 6 human corrections · 115 large tool outputs · 1 permission request · 62 repeated commands · 15 repeated failures · 36 repeated reads · 11 repeated skill loads · 63 sessions with usage bound to no task · 1 session without usage records

Details: snapshot, prices, coverage, counters
  • Task toolboxmd/chromeria#133: outcome unknown (recorded acceptance only; a finished process never implies it).
  • Proof: fix(server): preserve stored effort for agent-started turns #133
  • Snapshot 194dceead676383336ff932fae09cdf6f47d1ab960c5e09fd715b2cff102fc42, records up to 2026-10-02 07:33 UTC.
  • 64 sessions on claude, codex; AgentsMD 14.6.0.
  • Not counted: 3,590 responses (at least $45.42) in sessions shared with other PRs that worked in no single PR's checkout.
  • 3,350 responses in these sessions worked on other PRs and are counted there.
  • Totals reconcile with the measured sessions: yes. Evidence complete: no.
  • Prices: list-price estimate from T3 local rate table (path withheld) as of 2026-09-28, schedule 696aae45933d0a684a015d3f08cfb6f1aa2fef02e7d272b4689e3a692ea04fff. Unknown prices stay unknown, never zero.
    • T3 LiteLLM rate table when present; bundled schedule covers the rest.
  • Native session usage or worker ownership is unavailable.
  • Harness-reported cost: none reported.
  • Usage totals are not billing. Subscription spending is separate and is never posted as spend.
  • Crashed runs are counted separately: 0.

Token counters by model (native counter semantics; never added across semantics):

Selected rates (USD per million tokens). These rates value the report at the selected schedule date; they do not establish historical prices or subscription spending.

Model Input tier Input Cache read Cache write Other output Reasoning

Other output and reasoning are priced without double counting inclusive native output. Missing rates remain unknown.

Local measurement from native records; usage totals are not billing. Updated in place by agent-observer publish.

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire 13.5 KiB 13.5 KiB +15 B (+0.1%) 15.1 KiB ✅
Codex Thread snapshot wire 7.1 KiB 7.1 KiB +2 B (+0.0%) 7.3 KiB ✅
Codex Live turn WebSocket wire 6.4 KiB 6.4 KiB +13 B (+0.2%) 7.8 KiB ✅
Codex Live turn WebSocket decoded 56.2 KiB 56.2 KiB 0 B (0.0%) 66.4 KiB ✅
Codex Live turn messages 9 9 0 (0.0%) 21 ✅
Claude Total thread wire 13.5 KiB 13.5 KiB +6 B (+0.0%) 15.1 KiB ✅
Claude Thread snapshot wire 7.1 KiB 7.1 KiB −5 B (−0.1%) 7.3 KiB ✅
Claude Live turn WebSocket wire 6.4 KiB 6.4 KiB +11 B (+0.2%) 7.8 KiB ✅
Claude Live turn WebSocket decoded 57.0 KiB 57.0 KiB 0 B (0.0%) 66.4 KiB ✅
Claude Live turn messages 9 9 0 (0.0%) 21 ✅

Baseline: 86d3269 · PR result: a4fee31 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 114.0 KiB
  • Claude decoded thread snapshot: 114.7 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XS vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(server): preserve stored effort for agent-started turns

1 participant