Repository navigation
feat(wight): continue v2 threads within timer and quota limits - #188
Conversation
Agent work on this PREstimated cost unknown · 0 responses · 45 sessions · 7.5 h wall time
Flags: 3 human corrections · 88 large tool outputs · 99 repeated commands · 11 repeated failures · 103 repeated reads · 12 repeated skill loads · 44 sessions with usage bound to no task · 1 session without usage records Details: snapshot, prices, coverage, counters
Token counters by model (native counter semantics; never added across semantics): Selected rates (USD per million tokens). These rates value the report at the selected schedule date; they do not establish historical prices or subscription spending.
Other output and reasoning are priced without double counting inclusive native output. Missing rates remain unknown. Local measurement from native records; usage totals are not billing. Updated in place by |
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: unavailable · PR result: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
8ca4657 to
c5013d9
Compare
|
What: Independent review of PR #188 at Review scopeI inspected the exact 26-file #173 diff (1,570 additions, 4 deletions), its three-commit range-diff from the prior reviewed candidate, and the old-base to new-base overlap. Implementation hunks are unchanged; only runtime import and map/allowlist append context moved. The new base overlaps #173 in The D30 proof is the direct post-planning The modularity audit matches the handoff: all 11 upstream-owned changed paths have the listed counts and one primary owner, larger-hunk explanations are present, and the recorded fork check passed with 20 features. The feature map places Wight behavior in its own modules and uses shared upstream hooks. Proof inspectedThe exact-head composed batch records 140/141 passing across 12 files. Its sole failure is the unchanged symlink-settings watcher test, which hit its existing 2-second timeout at 2,055 ms. The isolated exact-head The exact-head handoff records a fresh server typecheck, runtime lint/format, and fork check. Web, contracts, and shared typechecks were reused from the prior reviewed head because their source and relevant manifests/compiler inputs are unchanged across the rebase; the recorded old-head to new-head input diff is empty. No full repository suite, build, live provider run, or integrated browser pass is claimed. Planner-owned UI evidence remains pending by the agreed deferral. Full commands and logs are in the exact-head proof handoff. Elon recordRequirements and who asked: The parent requested an exact-head review; #173 and the Objective require timer/quota-bounded Wight, D30 admission safety, and coexistence with the merged #172 runtime; the user requires modular fork-owned behavior. Reviewer: GPT-6-Luna, max reasoning, Codex harness in T3 Code. |
What: Wight mode keeps v2 threads working until their timer or provider-instance quota pauses new turns.
Why: The v2 engine has no Wight continuation, and stale automatic sends must never append a message, run, or queued turn.
So what: CI and exact-head independent review pass on the unchanged candidate; the parent verifies delivery, and the planner captures the deferred integrated UI evidence.
Closes #173.
Wight lives in fork modules, with a Context.Reference admission hook and production layer registration. It dispatches through the existing serialized thread service, preserves stored model/effort, and checks immutable projection identity inside the EventSink transaction. Automatic sends never acknowledge retirement tokens. Timer expiry lets the current turn finish; quota checks use fresh per-instance windows. Merged #169 supplies reset/retry coordination and
autoResumeLimitedThreads=true; Wight consumes it without early resume or duplicate state.Typed settings preserve v1 indefinite, timed and expired activations through unrelated writes. The existing persistence test retains every Prism/project/roundtrip assertion. Web/desktop expose the header toggle and provider quota setting; native mobile controls remain absent as in v1.
Candidate and proof
c5013d924ccab7d9ac63349919b22adbbe4536e2.36d5c0169fc042930aa36e41248c1292cec781d0(fork/v2, includes Port the Prism stream clock and stale-turn detector #172). Only three Port Wight mode #173 commits rebased; no dependency copies. The exact overlap is runtime registration, feature map and allowlist. Both Port the Prism stream clock and stale-turn detector #172 stream-clock/stale-turn modules and Wight remain registered; range-diff shows only context shifts.serverSettings.test.tsrerun had 63 pass, including that watcher, without source/assertion/timeout edits. This is not a clean 141-pass batch; the watcher's underlying cause remains unknown. Both logs and rationale are retained.uptimeload <8.c5013d924c. Parent verified newestreview/independentstatussuccess, creatorlukemaj. Earlier-head verdict is not reused.EventSink.commitCommand, with positive control and rollback of message/run/queued-effect/receipt/event writes. Tests exercise the real serialized caller and stored high effort with provider process execution disabled.UI evidence pending: user-approved deferral to planner's ONE integrated pass on
fork/v2after all ports land, including before/after evidence. Browser and dev servers remain disabled for this worker.CI retry on unchanged head
Attempt 1 of CI run 37844060523 cancelled Test Server 3 at the unchanged 10-minute job limit, causing aggregate Check to fail. Attempt 2 was explicitly authorized and started with
gh run rerun 37844060523 --failed; it completed successfully on unchanged headc5013d924ccab7d9ac63349919b22adbbe4536e2. Test Server 3 passed: job 21:26:31 to 21:31:23 UTC (292 seconds), dependency install 21:27:31 to 21:27:48 (17 seconds), test step 21:27:48 to 21:31:21 (213 seconds; Vitest 212.48 seconds), 102 files passed and 1548 tests passed / 1 existing skipped. Aggregate Check passed at 21:31:36 UTC. Live PR checks are all completed success or intentional skipped; no newly introduced skip, timeout increase or source change. Newest exact-head review status remains success bylukemaj.Comparison against the same successful shard on PR #184 head
57a59b4fa1d26abe842a12827e776428dc2f2b9f(run 37841082589, merged as current base):Raw timestamped logs show Ubuntu
azure.archive.ubuntu.compackage-download pauses, including 234.677 seconds beforefonts-wqy-zenheiand 56.73 seconds beforelibfreetype6; Chromium CDN download took approximately 3 seconds and FFmpeg approximately 1 second. Passing files continued until 21:13:58 before the 21:14:02 cancellation. These logs support exhausted setup budget, not a Wight test hang. Local symlink-watcher timeout remains separately documented above; its CI file passed in 2136 ms. No source change, timeout increase or test skip is included in #188; the parent owns the separate CI mirror-action concern. Existing exact-head independent review remains valid.Upstream edit audit
Fork behavior stays in new modules. Upstream edits are hooks, fields, registration, approved fixture migration and user guidance. Every allowlisted file retains exactly one primary feature-map owner.
apps/server/src/orchestration-v2/Orchestrator.tsapps/server/src/orchestration-v2/runtimeLayer.tsapps/server/src/serverSettings.test.tsapps/web/src/components/chat/ChatHeader.tsxapps/web/src/components/settings/ProviderInstanceCard.tsxdocs/user/thread-sidebar.mdpackages/contracts/src/index.tspackages/contracts/src/orchestrationV2.tspackages/contracts/src/providerInstance.tspackages/contracts/src/settings.tspackages/shared/src/serverSettings.tsHunks over approximately 15 lines including diff context:
Orchestrator.ts(19-line hunk): attaches the fork guard and calls the fork admission hook under existing serialization, converting refusal into the existing error. Behavior stays in fork modules; no upstream planner copied or restructured.orchestrationV2.ts(18-line hunk): server-only immutable admission identity on the message type, preserving client/internal variants. Wire schema unchanged.thread-sidebar.md(23-line hunk): 17 lines of user guidance for timers, quota, reset recovery and mobile availability.serverSettings.test.tschanges only the fixture (5 added/1 removed); original title and every Prism assertion remain.forkSettings.tshas zero final diff. Upstream-first assessment and #166 define the port boundary.Elon record
Requirements and who asked: Objective/#173 and parent require timer/quota-bounded Wight, saved settings/effort, retirement safety, D30 atomic proof, modular hooks and one reviewed PR.
Deleted: Early resume, duplicate D6' recovery ownership, wire idle mode, independent retirement state and unrelated formatter reflow.
Bottleneck: CI retry passed on the unchanged reviewed head; delivery verification remains with the parent and integrated UI evidence with the planner.
Checked myself: Failed-attempt setup/per-file comparison, successful attempt-2 raw log (102 files, 1548 tests passed / 1 existing skipped), all live checks and newest exact-head review creator, exact overlap/range-diff and own-only rebase, preserved runtime hooks, 140-pass/1-timeout batch plus 63-pass isolated rerun, fresh affected server/lint/format/fork proof, unchanged-input proof reuse, all 11 upstream counts/hunks, and parent exact-head review/cost records.
Model: GPT-6.1-Sol (medium reasoning). Harness: Codex in T3 Code.