Skip to content

feat(scheduled-tasks): adopt upstream scheduler with fork checks and v1 import - #193

Merged
lukemaj merged 31 commits into
fork/v2from
feat/174-scheduled-tasks-v2
Oct 9, 2026
Merged

lukemaj merged 31 commits into
fork/v2from
feat/174-scheduled-tasks-v2

Conversation

@lukemaj

@lukemaj lukemaj commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

What: Scheduled tasks run on upstream's single scheduler with the fork's outcome checks, one-shot and weekly triggers, Prism roles, shell command tasks, the Spectrum report fence and the Chromeria v1 task import.
Why: Upstream now ships its own scheduler; keeping the fork's v1 scheduler alongside it would duplicate tools and paths, and fork tasks would stop after the v1 cutover (#166).
So what: The CI repair passed focused checks, bounded exact-head review and cloud Lint/Test Web; human merge approval remains. After that, #176 wires its real Spectrum adapter to the exported port.

Closes #174.

Candidate: 7de85a13e1e78980064b7bf05fa2a8aa8b2def8b (tree 44b680c0fb88636fc7d20853824ae419521eb796), base 6eb69fe607c3940f20faa6ee9a8c4c9457e91af5. All 31 commits and 48 changed files are #174 scope; no dependency commits remain. #169/D35/D38 and #187 are merged. The mechanical #187 rebase preserved task source and test inputs; later server-only repairs have their own affected proof.

Problem

Chromeria v1 carried its own scheduler with check-gated tasks, one-shot and weekly schedules, Prism roles and shell command tasks. V2 adopts upstream's ScheduledTaskService, table, RPCs, clients and MCP tools, so these features had to move onto that one scheduler without forking it, and the existing v1 tasks had to come along.

How

Fork behavior lives in apps/server/src/scheduledTaskChecks/. Upstream files get only small hooks.

  • Dispatch: a ScheduledTaskDispatchPolicy Context.Reference is read inside upstream's runTask, before markRunning. It decides between upstream dispatch, a skip, or a fork dispatch, through dispatchVia. Upstream behavior is the default, and webhook tasks are unchanged.
  • Outcome checks: a check keeps the same thread working until its pinned check passes, with durable same-thread retries. A usage limit waits for reset. A retired thread, exhausted retries or a report that needs the user lead to needs-you, and Run now resumes it in its own thread. Sends are recorded with pinned payloads before dispatch, so a crash redelivers the same command once.
  • Triggers: once (an absolute instant with a durable single-fire identity) and weekly (weekdays, multiple times, IANA time zone, and minute spreading within a window). These are upstream schedule union members, read-only on clients and editable through MCP.
  • Prism roles: a task with an explicit role resolves and validates its launch through PrismService; Prism's own capacity routing may pick an eligible fallback. A post to a bound thread keeps that thread's stored model selection and takes only the role's kit; it never falls back to the task's modelSelection. A role that cannot be honored becomes a retry.
  • Command tasks: one shell command per run, in a detached process group with timeout and an output tail. Web and desktop alert when a task turns from passing to failing. The live task subscription leaves command output out; Settings fetches it on an explicit Show output through the existing list RPC, with a retry on failure, and never shows a stale result while fetching. The list and the agent tools keep the output.
  • Spectrum report fence (D40 planner decision, reviewed by Sol):
    • Completion requires every bound Spectrum to settle or retire. Its current report turn must complete through exact continuations, unless the user explicitly abandons that current report; abandonment never releases an active Spectrum. After a passing check, a waiting fence holds the run before rechecking. A failed check can record a retry, but its recorded send stays held until the fence releases.
    • A report that needs the user makes the run need the user, with the report's reason, before any check, retry, resend or redelivery.
    • The settle transaction re-reads the run's own work, through Port Prism as a service behind upstream delegation #169's merged recovery reader, and the fence.
    • Port Spectrum council and free discussions #176 provides the real Spectrum adapter through the exported ScheduledTaskSpectra port.
  • v1 import: a forkV1Backfills step imports the latest v1 scheduler.state-set per task. Imported tasks are disabled and their v1 runs stay inert history; a manual Run now starts a new run. Unsupported tasks keep their payload privately, and logs carry ids and bounded kinds only.
  • Deletion and start: deleting a task runs the judged-thread guard, upstream's delete and the fork state under the task's lock, in one transaction; the MCP tool passes its calling thread through a fork Context.Reference, so the lock never nests. An admitted send re-reads its run and the D40 fence under that lock before it goes out: a deleted task or a moved-on run sends nothing, a report that needs the user makes the run need the user, and a waiting report keeps the recorded send for redelivery. A command run captures its admitted command, finds its workspace, then re-checks ownership and spawns that captured command under the lock; the fork's runner returns at the spawn, so the lock is never held while a command runs. Existing in-memory execution bookkeeping gives each admission its own identity: a successful delete invalidates it, and an old command, result or deferred agent callback cannot act on a recreated task with reused ids. Failed deletes preserve it; cleanup removes only its own identity. Drive and restart recover the current persisted intent without requiring that ephemeral identity.
  • New-thread runs: a launch refused before its thread existed (a rejected receipt or a pre-shell failure) is launched again as a new send with new ids, on the same deterministic thread. Evidence is exact: no shell and no accepted receipt for any of the run's sends. A thread lost after any accepted send still needs the user.
  • Defects found and fixed:
    • The observer decoded stored runs with a type schema that cannot read ISO date strings. A landed send was therefore never seen as landed, and would have been redelivered on every tick. It now uses the published OrchestrationV2RunJson codec.
    • The fence and the launch evidence read orchestration_v2_command_receipts, a legacy table nothing writes. Both now read the live orchestration_command_receipts. Earlier fixtures seeded the legacy table and hid this; they now record through the real CommandReceiptStore.
  • Active Spectra: an active Spectrum never releases a run, but its current report's needs-you makes the run need the user. This is a contract-gap repair against the exposed port, not Port Spectrum council and free discussions #176 runtime coverage.

Upstream files (diff 6eb69fe607..7de85a13e1)

File Lines Largest hunk
apps/mobile/src/features/settings/SettingsScheduledTasksRouteScreen.tsx +9/-4 3
apps/mobile/src/features/settings/scheduledTaskDraft.ts +10/-2 3
apps/server/src/mcp/OrchestratorMcpService.ts +21/-6 7
apps/server/src/orchestration-v2/runtimeLayer.ts +12/-1 10
apps/server/src/persistence/forkV1Backfills.ts (fork file) +13/-0 9
apps/server/src/scheduledTasks/Schedule.ts +13/-1 8
apps/server/src/scheduledTasks/ScheduledTaskService.ts +19/-3 5
apps/web/src/components/ThreadNotificationCoordinator.badge.test.tsx +2/-0 2
apps/web/src/components/ThreadNotificationCoordinator.test.tsx +2/-0 2
apps/web/src/components/ThreadNotificationCoordinator.tsx +5/-1 5
apps/web/src/components/settings/ScheduledTasksSettings.tsx +8/-1 2
apps/web/src/components/settings/scheduledTasksSettings.logic.ts +14/-3 10
packages/client-runtime/package.json +4/-0 4
packages/client-runtime/src/state/server.ts +5/-0 5
packages/contracts/src/index.ts +1/-0 1
packages/contracts/src/orchestratorMcp.ts +4/-0 1
packages/contracts/src/scheduledTask.ts +4/-0 1
docs/user/project-settings.md +25/-0 25
docs/fork-features.md (fork map) +71/-0 71
scripts/fork-upstream-edits.txt (fork allowlist) +15/-0 15

Largest-hunk counts use complete Git --unified=3 hunks and count changed lines, so nearby edits merged by Git count together. Hunks over 15 lines:

  • docs/user/project-settings.md (25): the feature's new user section. AGENTS.md asks for one section per major feature, so this is a single section, not code.
  • docs/fork-features.md (71): the required scheduled-tasks feature map entry, listing new, upstream and shared files and the watch keywords. It is a fork-owned file.

Every upstream code hunk is 10 lines or fewer. Schedule.ts carries one hook inside upstream's unused describeSchedule, because it narrows the widened union for upstream's own branches below it.

Proof (listed final and revalidation steps passed fresh, fail-closed load gates below 8; mapped to 7de85a13e1)

The ownership and waiting-send repairs (2e7449d5b6, f481862c31, bb92caed3b, d8fea9e9fc, 6e9b42e876, 4629e8e808) touch only fork server files, so client, contracts, mobile and subset proof carries forward unchanged. Historical proof-focused-2 ran at load 22.49 in violation of the resource policy; its result is not resource-compliant evidence. The Issue preserves that violation and the subsequent fail-closed double-gate repair.

  • pnpm exec tsc --noEmit, all exit 0:

    • server on 4629e8e808;
    • client-runtime and mobile on 51ef1e3a55, unchanged after for what they check;
    • web on 011cac3aeb;
    • contracts on f0411f5d7d, unchanged since.
  • pnpm exec vp lint on the 41 changed TypeScript files: exit 0 on 8cc63787b7, then on each file changed since, on 011cac3aeb and f481862c31, with all three waiting-send files revalidated on bb92caed3b, both command-edit files revalidated on d8fea9e9fc, and all three admission-ownership files revalidated on 4629e8e808. The single warning predates this PR (upstream's unused describeSchedule).

  • pnpm exec vp test run apps/server/src/scheduledTaskChecks/ on 4629e8e808: 7 files, 75 passed, exit 0. It covers:

    • the engine, including deletion between admission and a send or a command spawn, the D40 fence at the send, deletion during workspace lookup, waiting-send recovery across restart with unchanged ids and payload, and command edits that affect only later runs, and delete/recreate aliases at spawn, result write, deferred send and cleanup;
    • deletion through the real service between admission and a send, a judged thread's delete racing its run's start, and a report that comes to need the user after Run now admits a send;
    • the real-orchestrator integration, including the rejected-launch relaunch across restarts with the deleted-thread guard, real-receipt release and hold, the output-free live subscription, the final-transaction race, the D40 lifecycle and needs-you before any check;
    • the fence on Port Spectrum council and free discussions #176's 7cb115d306 schema fixture, including the active-Spectrum boundary;
    • D38 history, schedules, the command runner and the synthetic v1 import.

    The actual-data subset test skips here by design.

  • Client-runtime and web scheduled-task tests on 8cc63787b7: 31 passed. Contracts and mobile tests on 6bedc0e7a5, inputs unchanged: 71 passed with the others.

  • ACTUAL-DATA SUBSET v1 import, standalone with its env vars, on 6bedc0e7a5. It is reused because the import, backfills and harness are unchanged. Result: 1 imported (agent:once), 0 unsupported, 1 bound thread, 1 inert run, v1 fields preserved; a rerun and a restart changed nothing. It is not a full-database snapshot.

  • Same-id ownership proof combines real schedule-tool/delete/Run now calls with a stable request key at a fixed TestClock (proving task/run ids really repeat), and faithful engine Deferred barriers (proving old commands, results and callbacks cannot cross deletion). The real command-service test is identity reachability evidence, not a service-level negative-spawn oracle. The condition is deliberately exercised; its production frequency is not claimed.

  • Each review regression failed on the old source first, on its behavior: the engine and real relaunch, the active Spectrum, both legacy-table receipt cases, the subscription output marker, the sends and the command after a delete, the send over needs-you and while waiting, the judged thread's racing delete, and the command spawned after a delete during its workspace lookup, and an admitted command changed during that lookup, and the command/result/agent-callback alias after recreation. The first agent-red batch also contained an unrelated missing-thread-worktree fixture failure; that is not regression evidence, and the corrected real-id case passed.

  • node scripts/fork-features.mjs check and bash scripts/fork-check.sh: OK on 011cac3aeb; the later commits add no file and touch no upstream file.

CI repair and current review

CI on 4629e8e808 failed: Lint reported 11 unused exports; Test Web reported 32 failures in two notification tests. The same failures reproduced locally before the repair: useAtomRefresh was missing from their atom mock, reached through the newly mounted scheduler notifier. This was fixture composition, not a demonstrated notification-policy defect.

7de85a13e1 privatizes all 11 same-file helpers/constants without removing the #176 port or changing runtime logic. Each existing notification test gets two fixture lines supplying query data:null (not loaded/no scheduler data); every assertion is unchanged. A separate fork-owned component test exercises actual command health transitions, alert dedupe, sound/focus, navigation and desktop badge handoff with controlled task data. It is not real RPC integration. The two small upstream test hooks are allowlisted.

Affected proof on the clean candidate: 35 notification/component tests passed; 75 server tests passed plus one actual-data env skip; targeted lint passed. Server/web typechecks and both exact CI knip invocations passed on clean 491abcdbf3, whose tree 44b680c0fb88636fc7d20853824ae419521eb796 is identical to the final 7de85a13e1 tree. No suppressions or waivers were added. The unchanged client/subset proof above remains reusable.

The previous 4629 source review is reused for unchanged code. The bounded independent repair review passed on the new exact head; authenticated lukemaj posted review/independent=success on 7de85a13e1e78980064b7bf05fa2a8aa8b2def8b. Cloud Lint and Test Web passed on 7de85a13e1: both exact knip checks are clean, and 491 web test files / 6,670 tests passed, including the existing 33 notification/badge tests and the 2 new notifier tests. Cloud Typecheck and aggregate Check passed too. The actual job logs were inspected; no waiver was used.

Full record: #174 (comment)

Cost measurement

AgentObserver refreshed [its summary] on the CI repair head(#193 (comment)). Publication succeeded, but attribution is incomplete: estimated cost is unknown, 50 sessions have usage bound to no task, and 1 has no usage records. The displayed 0 attributed responses is not zero work or a complete cost total. No manual ownership or cost was invented.

Limits

  • Port Spectrum council and free discussions #176's real Spectrum adapter and its integrated proof come after this merges. Until then no Spectra are bound, so the fence releases at once. The active-Spectrum repair is proved against the exposed port, not against Port Spectrum council and free discussions #176's runtime.
  • Command failure alerts are web and desktop only.
  • Browser/device UI evidence is deferred to the planner's integrated pass; no UI verification is claimed here. CI owns repo-wide checks, which were not run locally.
  • Accepted D40 risk: a report attempt that was partly processed may be processed twice when Spectrum retries.
  • Deleting a task does not stop a command process that already spawned; its result is not recorded.
  • All actual agent sends use the same locked ownership/fence guard, including redelivery and sends recorded during later reconcile steps. A waiting report holds the recorded intent until release; needs-you surfaces its reason.

Elon record

  • Requirements and who asked: the user's merge of scheduled tasks onto upstream's scheduler (Absorb upstream main (533 commits behind) #166, Merge scheduled tasks with upstream's scheduler #174); the D31–D40 decisions (D40 is the planner's decision, reviewed by Sol); the actual v1 import, as an actual-data subset; the review findings, confirmed by the coordinator and approved as routine fixes within Merge scheduled tasks with upstream's scheduler #174 and D40 scope.
  • Deleted: the fork's own v1 scheduler path, provider-turn delivery evidence and the D39 adapter audit, full-database copying (the data foundation's coverage), edits in upstream test files, the sends.length <= 1 launch guess, legacy receipt-table reads, command output on the live subscription, and the unlocked delete guard and fork-state forget.
  • Bottleneck: Human approval to merge the reviewed PR; cloud Lint/Test Web repairs are verified. The Port Prism as a service behind upstream delegation #169 recovery dependency is already merged.
  • Checked myself: clean exact-SHA proof logs with exit and load markers; the subset's byte identity against the source; the merged base tree equal to the reviewed dependency; upstream hunk sizes.

Implementation: Claude Opus 5.5 (claude-opus-5-5) in Claude Code through T3 Code. Independent review: GPT-6-Luna in Codex through Prism. Coordination: GPT-6.1-Sol in Codex through T3 Code.

🤖 Generated with Claude Code

lukemaj and others added 30 commits October 9, 2026 00:27
…s on upstream's scheduler

Server and contracts slice for #174. Upstream files get
small hooks only: one schedule-union spread, two read-model fields, the
dispatch policy and fire key in runTask, Schedule.ts guards, MCP tool hooks
and the runtime layer wrap. Client display and editors are not in this
checkpoint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ions, runner tests

In-place hooks: dispatchVia wraps upstream's launch/send calls and the MCP
schedule/update/delete wrappers wrap the existing service calls, so no
upstream expression is re-indented. Web and mobile show fork triggers
read-only through shared fork modules; command failures notify like thread
alerts. Client typecheck and the new runner tests have not run yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A role resolves through PrismService: a launch takes Prism's validated
model and kit, a post to a bound thread takes only the kit and keeps the
thread's model. With no eligible model the run retries; it never falls back
to the task's own model. Contracts use the merged PrismRole and PrismLane.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e (unwired)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ectra schema, D35 follower tests

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… settle transaction

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t needs-you makes the run need you), drop provider-turn evidence, schedule tests

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…chedule tests

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ing in the observer, integrated D40 and race proofs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-trip tests

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, Spectrum report wording

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hrough the real startup path

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…k, subset proof summary

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… values

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rface an active Spectrum's current report needs-you

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rd receipts through the real store

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…etches it from the list on demand

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r endless loading

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ching

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…judged thread's delete is guarded under the task's lock

The guard, upstream's delete and the fork state now run under the task's lock
in one transaction. The deferred send and command spawn re-read the state under
the lock and start nothing unless the task still owns the run; the deferred
send also re-reads the D40 fence there.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tarts

The workspace is found first; the ownership check and the real spawn then share
the task's lock, which is released once the process has spawned and never held
while it runs. The runner returns at the spawn, and its returned effect waits.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…to release

Drive's redelivery and its new sends now go out through the same locked
ownership and fence check as the deferred send, so a send held for a waiting
report stays unsent through later steps and restarts, then goes out once with
its recorded id and payload.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… made while it starts

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eated task's same-id run never borrows an older one

A successful delete ends its runs' admissions; a command fiber checks its exact
admission before its spawn and before writing its result, and its cleanup ends
only its own admission. Red and clean-SHA proof are still pending.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… of the same id

Deferred sends from a fire or Run now now carry their admission too, sharing
the command bookkeeping: only the exact admission sends, a successful delete
ends it, and a step recovers stored runs without one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lukemaj

lukemaj commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

What: Independent exact-head source review passed for 4629e8e8086defb61259d6622d105f561161687a.
Why: The final admission ownership changes close the confirmed delete/recreate alias without changing durable run identities.
So what: No blocking source findings remain; the PR awaits CI and human merge approval, with UI and #176 runtime integration explicitly deferred.

Verdict: pass for the frozen source

No actionable findings remain for 4629e8e, reviewed against base 6eb69fe. The base is its merge base; the candidate tree is 9a3b25bd2e769746105909bde0b7b63153c8f730. The frozen commit changes the engine and its two test files relative to d8.

The final ownership changes address the same-ID delete/recreate race: deletion clears admissions after the transaction commits while holding the task lock; command spawn rechecks the exact admission and current run under that lock after workspace lookup; results and cleanup require the same admission object. Deferred agent sends also recheck their admission and the current report fence under the lock. Drive can recover the current persisted intent without an ephemeral admission. Command ownership and cleanup · Deferred send guard · Fence and recovery path · Deletion invalidation

Cumulative finding history

  • Repaired: command capture across workspace lookup; serialized deletion and spawn/result ownership; same-ID command and deferred-send aliasing; D40 waiting and needs-you checks on dispatch/recovery; reads and fixtures now use the actual command receipt owner; and compact subscription output with demand reads and query error/pending handling.
  • Withdrawn: the earlier claim that Port Spectrum council and free discussions #176 had a reachable active-plus-current-needs-you runtime bug relied on a cross-Spectrum test, not that state. The store/schema permit the state, and the approved contract-gap repair covers it, but Port Spectrum council and free discussions #176 has no implemented report producer/adapter. Runtime reachability remains unknown, not a finding.
  • The final tests cover command and agent same-ID recreation and verify repeated IDs through the real schedule service. The tests and service-level ID cases support that review; the actual upstream dispatch consumer was also checked.

The writer and parent report 75 passing, 1 environment skip, server types 0 errors, lint 0 findings; I did not run checks. The reported old-source red proof included one real engine failure and one unrelated fixture failure; the corrected real-ID case passed. This review makes no browser/UI-pass claim and no full-database migration claim; no FoundationDB database is involved.

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Oct 9, 2026
@lukemaj

lukemaj commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor Author

Agent work on this PR

Estimated cost unknown · 0 responses · 51 sessions · 10.7 h wall time

Model Responses Tokens Estimated cost

Flags: 5 human corrections · 111 large tool outputs · 132 repeated commands · 11 repeated failures · 123 repeated reads · 18 repeated skill loads · 50 sessions with usage bound to no task · 1 session without usage records

Details: snapshot, prices, coverage, counters
  • Task toolboxmd/chromeria#174: outcome unknown (recorded acceptance only; a finished process never implies it).
  • Proof: Merge scheduled tasks with upstream's scheduler #174
  • Snapshot b439863e35fa5c2b039f2161cb00d7dfa2ee40188564d1fa2045ba69045fda6b, records up to 2026-10-09 00:21 UTC.
  • 51 sessions on claude, codex; AgentsMD 14.6.0.
  • Not counted: 6,863 responses (at least $98.69) in sessions shared with other PRs that worked in no single PR's checkout.
  • 6,917 responses in these sessions worked on other PRs and are counted there.
  • Totals reconcile with the measured sessions: yes. Evidence complete: no.
  • Prices: list-price estimate from T3 local rate table (path withheld) as of 2026-09-28, schedule 696aae45933d0a684a015d3f08cfb6f1aa2fef02e7d272b4689e3a692ea04fff. Unknown prices stay unknown, never zero.
    • T3 LiteLLM rate table when present; bundled schedule covers the rest.
  • Native session usage or worker ownership is unavailable.
  • Harness-reported cost: none reported.
  • Usage totals are not billing. Subscription spending is separate and is never posted as spend.
  • Crashed runs are counted separately: 0.

Token counters by model (native counter semantics; never added across semantics):

Selected rates (USD per million tokens). These rates value the report at the selected schedule date; they do not establish historical prices or subscription spending.

Model Input tier Input Cache read Cache write Other output Reasoning

Other output and reasoning are priced without double counting inclusive native output. Missing rates remain unknown.

Local measurement from native records; usage totals are not billing. Updated in place by agent-observer publish.

@github-actions

github-actions Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 5.0 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.8 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.9 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 5.0 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.8 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 21.2 KiB — 29.3 KiB ✅
Claude Live turn messages — 1 — 8 ✅

Baseline: unavailable · PR result: 7de85a1 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 108.5 KiB
  • Claude decoded thread snapshot: 108.8 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

…ification tests never load scheduler data

Knip found 11 exports used only in their own file; they are no longer exported.
The coordinator's two upstream test files mocked @effect/atom-react for its
shell reads only, so the scheduled command notifier it now mounts crashed on the
missing useAtomRefresh. They now mock the environment query to return no data
(the scheduler's list never loads, not a loaded empty list), every assertion
unchanged, and a fork test drives the notifier with real task data.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lukemaj

lukemaj commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

What: Bounded independent source review passed on the CI repair head 7de85a13e1e78980064b7bf05fa2a8aa8b2def8b.
Why: The repair preserves every notification assertion and makes same-file helpers private without removing the Spectrum port.
So what: Source review is complete; cloud Lint and Test Web still need confirmation before human merge approval.

The existing fixtures return query data:null, meaning not loaded/no scheduler data, not a loaded empty task list. The reviewer's verdict is relayed unchanged below.

Verdict: pass

No actionable findings for 7de85a1, reviewed against base 6eb69fe. I verified its tree 44b680c0fb88636fc7d20853824ae419521eb796 matches 491 exactly, with an empty diff; the merge base is the requested base. Frozen commit · 4629-to-7de comparison

The CI repair stays within scope: it privatizes the 11 internal exports without removing handoff ports, gives the two existing coordinator tests an empty scheduler-query mock without changing their assertions, and adds component-level notification coverage for failure transitions, sound, focus, navigation, and handoff to the badge owner. The new test does not exercise real RPC or assert integrated badge-count behavior. New notification test

Writer and parent report the targeted notification tests passing (35), server tests passing (75, with 1 environment skip), and type, lint, and knip checks clean. I did not rerun proof; the affected cloud jobs are still running. The earlier cumulative review stands, including the withdrawn #176 runtime claim and its unresolved producer reachability; this review adds no new findings.

@lukemaj
lukemaj merged commit 2589b89 into fork/v2 Oct 9, 2026
28 of 30 checks passed
@lukemaj
lukemaj deleted the feat/174-scheduled-tasks-v2 branch October 9, 2026 00:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant