Skip to content

feat(wight): continue v2 threads within timer and quota limits - #188

Merged
lukemaj merged 3 commits into
fork/v2from
feat/173-wight-v2
Oct 8, 2026
Merged

lukemaj merged 3 commits into
fork/v2from
feat/173-wight-v2

Conversation

@lukemaj

@lukemaj lukemaj commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

What: Wight mode keeps v2 threads working until their timer or provider-instance quota pauses new turns.
Why: The v2 engine has no Wight continuation, and stale automatic sends must never append a message, run, or queued turn.
So what: CI and exact-head independent review pass on the unchanged candidate; the parent verifies delivery, and the planner captures the deferred integrated UI evidence.

Closes #173.

Wight lives in fork modules, with a Context.Reference admission hook and production layer registration. It dispatches through the existing serialized thread service, preserves stored model/effort, and checks immutable projection identity inside the EventSink transaction. Automatic sends never acknowledge retirement tokens. Timer expiry lets the current turn finish; quota checks use fresh per-instance windows. Merged #169 supplies reset/retry coordination and autoResumeLimitedThreads=true; Wight consumes it without early resume or duplicate state.

Typed settings preserve v1 indefinite, timed and expired activations through unrelated writes. The existing persistence test retains every Prism/project/roundtrip assertion. Web/desktop expose the header toggle and provider quota setting; native mobile controls remain absent as in v1.

Candidate and proof

  • Frozen head: c5013d924ccab7d9ac63349919b22adbbe4536e2.
  • Base: 36d5c0169fc042930aa36e41248c1292cec781d0 (fork/v2, includes Port the Prism stream clock and stale-turn detector #172). Only three Port Wight mode #173 commits rebased; no dependency copies. The exact overlap is runtime registration, feature map and allowlist. Both Port the Prism stream clock and stale-turn detector #172 stream-clock/stale-turn modules and Wight remain registered; range-diff shows only context shifts.
  • Current exact-head proof and commands: composed Wight/settings, recovery/Stop/retirement and Port the Prism stream clock and stale-turn detector #172 integration batch had 140 pass / 1 unchanged watcher timeout. Isolated serverSettings.test.ts rerun had 63 pass, including that watcher, without source/assertion/timeout edits. This is not a clean 141-pass batch; the watcher's underlying cause remains unknown. Both logs and rationale are retained.
  • Fresh affected server typecheck, runtime lint/format and fork validation pass. Unaffected web/contracts/shared package proof and unchanged-file lint/format are reused: their sources, contracts, manifests, lockfile and compiler inputs are unchanged; exact input diff is empty and environment/dependencies are identical. Commands and reuse rationale are in the handoff. Heavy batches were serial with immediate uptime load <8.
  • Independent new-head review passes on exact c5013d924c. Parent verified newest review/independent status success, creator lukemaj. Earlier-head verdict is not reused.
  • D30 proof distinguishes stale admission snapshots from post-planning EventSink.commitCommand, with positive control and rollback of message/run/queued-effect/receipt/event writes. Tests exercise the real serialized caller and stored high effort with provider process execution disabled.
  • Earlier failed-before rationale: receipt oracle corrected to transactional rollback, copied run ordinal and introduced type/directive errors fixed; no assertions or proof waived.
  • Agent Observer cost attribution was refreshed by the parent after this push. First sync exited 130 at high load; subsequent below-8 retry succeeded and updated the same comment. Attribution remains partial and estimated cost unknown; no numeric cost claim.
  • Scoped unchanged-file lint has only pre-existing warnings. No repo-wide checks or live provider/UI verification.

UI evidence pending: user-approved deferral to planner's ONE integrated pass on fork/v2 after all ports land, including before/after evidence. Browser and dev servers remain disabled for this worker.

CI retry on unchanged head

Attempt 1 of CI run 37844060523 cancelled Test Server 3 at the unchanged 10-minute job limit, causing aggregate Check to fail. Attempt 2 was explicitly authorized and started with gh run rerun 37844060523 --failed; it completed successfully on unchanged head c5013d924ccab7d9ac63349919b22adbbe4536e2. Test Server 3 passed: job 21:26:31 to 21:31:23 UTC (292 seconds), dependency install 21:27:31 to 21:27:48 (17 seconds), test step 21:27:48 to 21:31:21 (213 seconds; Vitest 212.48 seconds), 102 files passed and 1548 tests passed / 1 existing skipped. Aggregate Check passed at 21:31:36 UTC. Live PR checks are all completed success or intentional skipped; no newly introduced skip, timeout increase or source change. Newest exact-head review status remains success by lukemaj.

Comparison against the same successful shard on PR #184 head 57a59b4fa1d26abe842a12827e776428dc2f2b9f (run 37841082589, merged as current base):

Observation #188 attempt 1 #184 base comparator
Chromium dependency-install step 407 seconds 17 seconds
Test-step wall time 134 seconds before cancellation 207 seconds, successful
ProviderSwitch file 37443 ms 44396 ms
Settings file 2136 ms 3444 ms
Wight admission 632 ms, 9 tests pass Absent from base
Completed test files 72 before cutoff 101

Raw timestamped logs show Ubuntu azure.archive.ubuntu.com package-download pauses, including 234.677 seconds before fonts-wqy-zenhei and 56.73 seconds before libfreetype6; Chromium CDN download took approximately 3 seconds and FFmpeg approximately 1 second. Passing files continued until 21:13:58 before the 21:14:02 cancellation. These logs support exhausted setup budget, not a Wight test hang. Local symlink-watcher timeout remains separately documented above; its CI file passed in 2136 ms. No source change, timeout increase or test skip is included in #188; the parent owns the separate CI mirror-action concern. Existing exact-head independent review remains valid.

Upstream edit audit

Fork behavior stays in new modules. Upstream edits are hooks, fields, registration, approved fixture migration and user guidance. Every allowlisted file retains exactly one primary feature-map owner.

Upstream-owned file Added Removed
apps/server/src/orchestration-v2/Orchestrator.ts 16 1
apps/server/src/orchestration-v2/runtimeLayer.ts 10 1
apps/server/src/serverSettings.test.ts 5 1
apps/web/src/components/chat/ChatHeader.tsx 4 0
apps/web/src/components/settings/ProviderInstanceCard.tsx 11 0
docs/user/thread-sidebar.md 17 0
packages/contracts/src/index.ts 1 0
packages/contracts/src/orchestrationV2.ts 12 1
packages/contracts/src/providerInstance.ts 2 0
packages/contracts/src/settings.ts 6 0
packages/shared/src/serverSettings.ts 4 0

Hunks over approximately 15 lines including diff context:

  • Orchestrator.ts (19-line hunk): attaches the fork guard and calls the fork admission hook under existing serialization, converting refusal into the existing error. Behavior stays in fork modules; no upstream planner copied or restructured.
  • orchestrationV2.ts (18-line hunk): server-only immutable admission identity on the message type, preserving client/internal variants. Wire schema unchanged.
  • thread-sidebar.md (23-line hunk): 17 lines of user guidance for timers, quota, reset recovery and mobile availability.

serverSettings.test.ts changes only the fixture (5 added/1 removed); original title and every Prism assertion remain. forkSettings.ts has zero final diff. Upstream-first assessment and #166 define the port boundary.

Elon record

Requirements and who asked: Objective/#173 and parent require timer/quota-bounded Wight, saved settings/effort, retirement safety, D30 atomic proof, modular hooks and one reviewed PR.
Deleted: Early resume, duplicate D6' recovery ownership, wire idle mode, independent retirement state and unrelated formatter reflow.
Bottleneck: CI retry passed on the unchanged reviewed head; delivery verification remains with the parent and integrated UI evidence with the planner.
Checked myself: Failed-attempt setup/per-file comparison, successful attempt-2 raw log (102 files, 1548 tests passed / 1 existing skipped), all live checks and newest exact-head review creator, exact overlap/range-diff and own-only rebase, preserved runtime hooks, 140-pass/1-timeout batch plus 63-pass isolated rerun, fresh affected server/lint/format/fork proof, unchanged-input proof reuse, all 11 upstream counts/hunks, and parent exact-head review/cost records.

Model: GPT-6.1-Sol (medium reasoning). Harness: Codex in T3 Code.

@lukemaj lukemaj mentioned this pull request Oct 8, 2026
@lukemaj

lukemaj commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor Author

Agent work on this PR

Estimated cost unknown · 0 responses · 45 sessions · 7.5 h wall time

Model Responses Tokens Estimated cost

Flags: 3 human corrections · 88 large tool outputs · 99 repeated commands · 11 repeated failures · 103 repeated reads · 12 repeated skill loads · 44 sessions with usage bound to no task · 1 session without usage records

Details: snapshot, prices, coverage, counters
  • Task toolboxmd/chromeria#173: outcome unknown (recorded acceptance only; a finished process never implies it).
  • Proof: Port Wight mode #173
  • Snapshot c5dfbed344b16114e2a98fecc58b337efa6cfed37cbc78fbc1800de6827e9683, records up to 2026-10-08 21:07 UTC.
  • 45 sessions on claude, codex; AgentsMD 14.6.0.
  • Not counted: 5,738 responses (at least $124.88) in sessions shared with other PRs that worked in no single PR's checkout.
  • 4,879 responses in these sessions worked on other PRs and are counted there.
  • Totals reconcile with the measured sessions: yes. Evidence complete: no.
  • Prices: list-price estimate from T3 local rate table (path withheld) as of 2026-09-28, schedule 696aae45933d0a684a015d3f08cfb6f1aa2fef02e7d272b4689e3a692ea04fff. Unknown prices stay unknown, never zero.
    • T3 LiteLLM rate table when present; bundled schedule covers the rest.
  • Native session usage or worker ownership is unavailable.
  • Harness-reported cost: none reported.
  • Usage totals are not billing. Subscription spending is separate and is never posted as spend.
  • Crashed runs are counted separately: 0.

Token counters by model (native counter semantics; never added across semantics):

Selected rates (USD per million tokens). These rates value the report at the selected schedule date; they do not establish historical prices or subscription spending.

Model Input tier Input Cache read Cache write Other output Reasoning

Other output and reasoning are priced without double counting inclusive native output. Missing rates remain unknown.

Local measurement from native records; usage totals are not billing. Updated in place by agent-observer publish.

@github-actions github-actions Bot added size:XL vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. labels Oct 8, 2026
@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 5.0 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.8 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.9 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 5.0 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.8 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 21.2 KiB — 29.3 KiB ✅
Claude Live turn messages — 2 — 8 ✅

Baseline: unavailable · PR result: c5013d9 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 108.5 KiB
  • Claude decoded thread snapshot: 108.8 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@lukemaj
lukemaj force-pushed the feat/173-wight-v2 branch from 8ca4657 to c5013d9 Compare October 8, 2026 21:03
@lukemaj

lukemaj commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

What: Independent review of PR #188 at c5013d924ccab7d9ac63349919b22adbbe4536e2 against 36d5c0169fc042930aa36e41248c1292cec781d0 found no blocking code issue in the reviewed scope.
Why: Rebase onto merged #172 overlaps runtime registration and feature ownership metadata; the exact diff retains both Prism stream-clock/stale-turn behavior and Wight admission/runtime layers.
So what: The PR author can refresh the same PR body with this candidate; the planner's integrated UI pass remains pending. This is not merge approval.

Review scope

I inspected the exact 26-file #173 diff (1,570 additions, 4 deletions), its three-commit range-diff from the prior reviewed candidate, and the old-base to new-base overlap. Implementation hunks are unchanged; only runtime import and map/allowlist append context moved. The new base overlaps #173 in runtimeLayer.ts, docs/fork-features.md, and scripts/fork-upstream-edits.txt. The resulting runtime retains PrismProviderEventIngestor, stream-clock provision, PrismStaleTurnMonitor, WightMode, and WightAdmission together. No dependency copy appears in the diff.

The D30 proof is the direct post-planning EventSink.commitCommand case: after a real Wight guard/plan and event set are prepared, the stale variant changes the projection before commit and rejects without message, run, queued work, receipt, outbox, or stored event writes. Its stable positive control commits the message, queued run, and outbox effect. The surrounding source still preserves D28 propagated-Stop checks, does not let server messages acknowledge retirement, and keeps saved model/effort. The v1 fixture contains indefinite, future-expiry, and expired Wight shapes; existing Prism/project/roundtrip assertions remain.

The modularity audit matches the handoff: all 11 upstream-owned changed paths have the listed counts and one primary owner, larger-hunk explanations are present, and the recorded fork check passed with 20 features. The feature map places Wight behavior in its own modules and uses shared upstream hooks.

Proof inspected

The exact-head composed batch records 140/141 passing across 12 files. Its sole failure is the unchanged symlink-settings watcher test, which hit its existing 2-second timeout at 2,055 ms. The isolated exact-head serverSettings.test.ts run then passed 63/63, including that watcher. The existing watcher timeout and assertions were not changed. I retain the combined-batch failure as a failure; its underlying cause is unknown, and this is not represented as one clean 141-pass batch. All Wight, recovery, Stop/retirement, and #172 integration cases passed in the composed run, including the D30 controls.

The exact-head handoff records a fresh server typecheck, runtime lint/format, and fork check. Web, contracts, and shared typechecks were reused from the prior reviewed head because their source and relevant manifests/compiler inputs are unchanged across the rebase; the recorded old-head to new-head input diff is empty. No full repository suite, build, live provider run, or integrated browser pass is claimed. Planner-owned UI evidence remains pending by the agreed deferral. Full commands and logs are in the exact-head proof handoff.

Elon record

Requirements and who asked: The parent requested an exact-head review; #173 and the Objective require timer/quota-bounded Wight, D30 admission safety, and coexistence with the merged #172 runtime; the user requires modular fork-owned behavior.
Deleted: Early resume, duplicate retirement/recovery ownership, and settings/provider I/O from the D30 guard remain excluded; no test timeout or assertion was relaxed.
Bottleneck: The sole PR author refreshes the current PR body after review; the planner's integrated UI pass remains pending.
Checked myself: Live PR base/head and authenticated reviewer; complete #173 diff and range-diff; #172 overlap and runtime composition; D30 guard/positive-control source; feature-map and allowlist ownership/counts; exact-head logs and proof-reuse inputs.

Reviewer: GPT-6-Luna, max reasoning, Codex harness in T3 Code.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant