Repository navigation
feat(scheduled-tasks): adopt upstream scheduler with fork checks and v1 import - #193
Conversation
…s on upstream's scheduler Server and contracts slice for #174. Upstream files get small hooks only: one schedule-union spread, two read-model fields, the dispatch policy and fire key in runTask, Schedule.ts guards, MCP tool hooks and the runtime layer wrap. Client display and editors are not in this checkpoint. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ions, runner tests In-place hooks: dispatchVia wraps upstream's launch/send calls and the MCP schedule/update/delete wrappers wrap the existing service calls, so no upstream expression is re-indented. Web and mobile show fork triggers read-only through shared fork modules; command failures notify like thread alerts. Client typecheck and the new runner tests have not run yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A role resolves through PrismService: a launch takes Prism's validated model and kit, a post to a bound thread takes only the kit and keeps the thread's model. With no eligible model the run retries; it never falls back to the task's own model. Contracts use the merged PrismRole and PrismLane. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e (unwired) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ectra schema, D35 follower tests Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… settle transaction Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t needs-you makes the run need you), drop provider-turn evidence, schedule tests Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…chedule tests Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ing in the observer, integrated D40 and race proofs Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-trip tests Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, Spectrum report wording Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hrough the real startup path Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…k, subset proof summary Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… values Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rface an active Spectrum's current report needs-you Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rd receipts through the real store Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…etches it from the list on demand Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r endless loading Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ching Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…judged thread's delete is guarded under the task's lock The guard, upstream's delete and the fork state now run under the task's lock in one transaction. The deferred send and command spawn re-read the state under the lock and start nothing unless the task still owns the run; the deferred send also re-reads the D40 fence there. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tarts The workspace is found first; the ownership check and the real spawn then share the task's lock, which is released once the process has spawned and never held while it runs. The runner returns at the spawn, and its returned effect waits. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…to release Drive's redelivery and its new sends now go out through the same locked ownership and fence check as the deferred send, so a send held for a waiting report stays unsent through later steps and restarts, then goes out once with its recorded id and payload. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… made while it starts Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eated task's same-id run never borrows an older one A successful delete ends its runs' admissions; a command fiber checks its exact admission before its spawn and before writing its result, and its cleanup ends only its own admission. Red and clean-SHA proof are still pending. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… of the same id Deferred sends from a fire or Run now now carry their admission too, sharing the command bookkeeping: only the exact admission sends, a successful delete ends it, and a step recovers stored runs without one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
What: Independent exact-head source review passed for Verdict: pass for the frozen sourceNo actionable findings remain for 4629e8e, reviewed against base 6eb69fe. The base is its merge base; the candidate tree is The final ownership changes address the same-ID delete/recreate race: deletion clears admissions after the transaction commits while holding the task lock; command spawn rechecks the exact admission and current run under that lock after workspace lookup; results and cleanup require the same admission object. Deferred agent sends also recheck their admission and the current report fence under the lock. Drive can recover the current persisted intent without an ephemeral admission. Command ownership and cleanup · Deferred send guard · Fence and recovery path · Deletion invalidation Cumulative finding history
The writer and parent report 75 passing, 1 environment skip, server types 0 errors, lint 0 findings; I did not run checks. The reported old-source red proof included one real engine failure and one unrelated fixture failure; the corrected real-ID case passed. This review makes no browser/UI-pass claim and no full-database migration claim; no FoundationDB database is involved. |
Agent work on this PREstimated cost unknown · 0 responses · 51 sessions · 10.7 h wall time
Flags: 5 human corrections · 111 large tool outputs · 132 repeated commands · 11 repeated failures · 123 repeated reads · 18 repeated skill loads · 50 sessions with usage bound to no task · 1 session without usage records Details: snapshot, prices, coverage, counters
Token counters by model (native counter semantics; never added across semantics): Selected rates (USD per million tokens). These rates value the report at the selected schedule date; they do not establish historical prices or subscription spending.
Other output and reasoning are priced without double counting inclusive native output. Missing rates remain unknown. Local measurement from native records; usage totals are not billing. Updated in place by |
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: unavailable · PR result: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
…ification tests never load scheduler data Knip found 11 exports used only in their own file; they are no longer exported. The coordinator's two upstream test files mocked @effect/atom-react for its shell reads only, so the scheduled command notifier it now mounts crashed on the missing useAtomRefresh. They now mock the environment query to return no data (the scheduler's list never loads, not a loaded empty list), every assertion unchanged, and a fork test drives the notifier with real task data. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
What: Bounded independent source review passed on the CI repair head The existing fixtures return query Verdict: passNo actionable findings for 7de85a1, reviewed against base 6eb69fe. I verified its tree The CI repair stays within scope: it privatizes the 11 internal exports without removing handoff ports, gives the two existing coordinator tests an empty scheduler-query mock without changing their assertions, and adds component-level notification coverage for failure transitions, sound, focus, navigation, and handoff to the badge owner. The new test does not exercise real RPC or assert integrated badge-count behavior. New notification test Writer and parent report the targeted notification tests passing (35), server tests passing (75, with 1 environment skip), and type, lint, and knip checks clean. I did not rerun proof; the affected cloud jobs are still running. The earlier cumulative review stands, including the withdrawn #176 runtime claim and its unresolved producer reachability; this review adds no new findings. |
What: Scheduled tasks run on upstream's single scheduler with the fork's outcome checks, one-shot and weekly triggers, Prism roles, shell command tasks, the Spectrum report fence and the Chromeria v1 task import.
Why: Upstream now ships its own scheduler; keeping the fork's v1 scheduler alongside it would duplicate tools and paths, and fork tasks would stop after the v1 cutover (#166).
So what: The CI repair passed focused checks, bounded exact-head review and cloud Lint/Test Web; human merge approval remains. After that, #176 wires its real Spectrum adapter to the exported port.
Closes #174.
Candidate:
7de85a13e1e78980064b7bf05fa2a8aa8b2def8b(tree44b680c0fb88636fc7d20853824ae419521eb796), base6eb69fe607c3940f20faa6ee9a8c4c9457e91af5. All 31 commits and 48 changed files are #174 scope; no dependency commits remain. #169/D35/D38 and #187 are merged. The mechanical #187 rebase preserved task source and test inputs; later server-only repairs have their own affected proof.Problem
Chromeria v1 carried its own scheduler with check-gated tasks, one-shot and weekly schedules, Prism roles and shell command tasks. V2 adopts upstream's
ScheduledTaskService, table, RPCs, clients and MCP tools, so these features had to move onto that one scheduler without forking it, and the existing v1 tasks had to come along.How
Fork behavior lives in
apps/server/src/scheduledTaskChecks/. Upstream files get only small hooks.ScheduledTaskDispatchPolicyContext.Referenceis read inside upstream'srunTask, beforemarkRunning. It decides between upstream dispatch, a skip, or a fork dispatch, throughdispatchVia. Upstream behavior is the default, and webhook tasks are unchanged.once(an absolute instant with a durable single-fire identity) andweekly(weekdays, multiple times, IANA time zone, and minute spreading within a window). These are upstream schedule union members, read-only on clients and editable through MCP.PrismService; Prism's own capacity routing may pick an eligible fallback. A post to a bound thread keeps that thread's stored model selection and takes only the role's kit; it never falls back to the task'smodelSelection. A role that cannot be honored becomes a retry.ScheduledTaskSpectraport.forkV1Backfillsstep imports the latest v1scheduler.state-setper task. Imported tasks are disabled and their v1 runs stay inert history; a manual Run now starts a new run. Unsupported tasks keep their payload privately, and logs carry ids and bounded kinds only.Context.Reference, so the lock never nests. An admitted send re-reads its run and the D40 fence under that lock before it goes out: a deleted task or a moved-on run sends nothing, a report that needs the user makes the run need the user, and a waiting report keeps the recorded send for redelivery. A command run captures its admitted command, finds its workspace, then re-checks ownership and spawns that captured command under the lock; the fork's runner returns at the spawn, so the lock is never held while a command runs. Existing in-memory execution bookkeeping gives each admission its own identity: a successful delete invalidates it, and an old command, result or deferred agent callback cannot act on a recreated task with reused ids. Failed deletes preserve it; cleanup removes only its own identity. Drive and restart recover the current persisted intent without requiring that ephemeral identity.OrchestrationV2RunJsoncodec.orchestration_v2_command_receipts, a legacy table nothing writes. Both now read the liveorchestration_command_receipts. Earlier fixtures seeded the legacy table and hid this; they now record through the realCommandReceiptStore.Upstream files (diff
6eb69fe607..7de85a13e1)apps/mobile/src/features/settings/SettingsScheduledTasksRouteScreen.tsxapps/mobile/src/features/settings/scheduledTaskDraft.tsapps/server/src/mcp/OrchestratorMcpService.tsapps/server/src/orchestration-v2/runtimeLayer.tsapps/server/src/persistence/forkV1Backfills.ts(fork file)apps/server/src/scheduledTasks/Schedule.tsapps/server/src/scheduledTasks/ScheduledTaskService.tsapps/web/src/components/ThreadNotificationCoordinator.badge.test.tsxapps/web/src/components/ThreadNotificationCoordinator.test.tsxapps/web/src/components/ThreadNotificationCoordinator.tsxapps/web/src/components/settings/ScheduledTasksSettings.tsxapps/web/src/components/settings/scheduledTasksSettings.logic.tspackages/client-runtime/package.jsonpackages/client-runtime/src/state/server.tspackages/contracts/src/index.tspackages/contracts/src/orchestratorMcp.tspackages/contracts/src/scheduledTask.tsdocs/user/project-settings.mddocs/fork-features.md(fork map)scripts/fork-upstream-edits.txt(fork allowlist)Largest-hunk counts use complete Git
--unified=3hunks and count changed lines, so nearby edits merged by Git count together. Hunks over 15 lines:docs/user/project-settings.md(25): the feature's new user section. AGENTS.md asks for one section per major feature, so this is a single section, not code.docs/fork-features.md(71): the requiredscheduled-tasksfeature map entry, listing new, upstream and shared files and the watch keywords. It is a fork-owned file.Every upstream code hunk is 10 lines or fewer.
Schedule.tscarries one hook inside upstream's unuseddescribeSchedule, because it narrows the widened union for upstream's own branches below it.Proof (listed final and revalidation steps passed fresh, fail-closed load gates below 8; mapped to
7de85a13e1)The ownership and waiting-send repairs (
2e7449d5b6,f481862c31,bb92caed3b,d8fea9e9fc,6e9b42e876,4629e8e808) touch only fork server files, so client, contracts, mobile and subset proof carries forward unchanged. Historicalproof-focused-2ran at load 22.49 in violation of the resource policy; its result is not resource-compliant evidence. The Issue preserves that violation and the subsequent fail-closed double-gate repair.pnpm exec tsc --noEmit, all exit 0:4629e8e808;51ef1e3a55, unchanged after for what they check;011cac3aeb;f0411f5d7d, unchanged since.pnpm exec vp linton the 41 changed TypeScript files: exit 0 on8cc63787b7, then on each file changed since, on011cac3aebandf481862c31, with all three waiting-send files revalidated onbb92caed3b, both command-edit files revalidated ond8fea9e9fc, and all three admission-ownership files revalidated on4629e8e808. The single warning predates this PR (upstream's unuseddescribeSchedule).pnpm exec vp test run apps/server/src/scheduledTaskChecks/on4629e8e808: 7 files, 75 passed, exit 0. It covers:7cb115d306schema fixture, including the active-Spectrum boundary;The actual-data subset test skips here by design.
Client-runtime and web scheduled-task tests on
8cc63787b7: 31 passed. Contracts and mobile tests on6bedc0e7a5, inputs unchanged: 71 passed with the others.ACTUAL-DATA SUBSET v1 import, standalone with its env vars, on
6bedc0e7a5. It is reused because the import, backfills and harness are unchanged. Result: 1 imported (agent:once), 0 unsupported, 1 bound thread, 1 inert run, v1 fields preserved; a rerun and a restart changed nothing. It is not a full-database snapshot.Same-id ownership proof combines real schedule-tool/delete/Run now calls with a stable request key at a fixed TestClock (proving task/run ids really repeat), and faithful engine Deferred barriers (proving old commands, results and callbacks cannot cross deletion). The real command-service test is identity reachability evidence, not a service-level negative-spawn oracle. The condition is deliberately exercised; its production frequency is not claimed.
Each review regression failed on the old source first, on its behavior: the engine and real relaunch, the active Spectrum, both legacy-table receipt cases, the subscription output marker, the sends and the command after a delete, the send over needs-you and while waiting, the judged thread's racing delete, and the command spawned after a delete during its workspace lookup, and an admitted command changed during that lookup, and the command/result/agent-callback alias after recreation. The first agent-red batch also contained an unrelated missing-thread-worktree fixture failure; that is not regression evidence, and the corrected real-id case passed.
node scripts/fork-features.mjs checkandbash scripts/fork-check.sh: OK on011cac3aeb; the later commits add no file and touch no upstream file.CI repair and current review
CI on
4629e8e808failed: Lint reported 11 unused exports; Test Web reported 32 failures in two notification tests. The same failures reproduced locally before the repair:useAtomRefreshwas missing from their atom mock, reached through the newly mounted scheduler notifier. This was fixture composition, not a demonstrated notification-policy defect.7de85a13e1privatizes all 11 same-file helpers/constants without removing the #176 port or changing runtime logic. Each existing notification test gets two fixture lines supplying querydata:null(not loaded/no scheduler data); every assertion is unchanged. A separate fork-owned component test exercises actual command health transitions, alert dedupe, sound/focus, navigation and desktop badge handoff with controlled task data. It is not real RPC integration. The two small upstream test hooks are allowlisted.Affected proof on the clean candidate: 35 notification/component tests passed; 75 server tests passed plus one actual-data env skip; targeted lint passed. Server/web typechecks and both exact CI knip invocations passed on clean
491abcdbf3, whose tree44b680c0fb88636fc7d20853824ae419521eb796is identical to the final7de85a13e1tree. No suppressions or waivers were added. The unchanged client/subset proof above remains reusable.The previous 4629 source review is reused for unchanged code. The bounded independent repair review passed on the new exact head; authenticated
lukemajpostedreview/independent=successon7de85a13e1e78980064b7bf05fa2a8aa8b2def8b. Cloud Lint and Test Web passed on7de85a13e1: both exact knip checks are clean, and 491 web test files / 6,670 tests passed, including the existing 33 notification/badge tests and the 2 new notifier tests. Cloud Typecheck and aggregate Check passed too. The actual job logs were inspected; no waiver was used.Full record: #174 (comment)
Cost measurement
AgentObserver refreshed [its summary] on the CI repair head(#193 (comment)). Publication succeeded, but attribution is incomplete: estimated cost is unknown, 50 sessions have usage bound to no task, and 1 has no usage records. The displayed 0 attributed responses is not zero work or a complete cost total. No manual ownership or cost was invented.
Limits
Elon record
sends.length <= 1launch guess, legacy receipt-table reads, command output on the live subscription, and the unlocked delete guard and fork-state forget.Implementation: Claude Opus 5.5 (
claude-opus-5-5) in Claude Code through T3 Code. Independent review: GPT-6-Luna in Codex through Prism. Coordination: GPT-6.1-Sol in Codex through T3 Code.🤖 Generated with Claude Code