Skip to content

feat(server): threads resume on their own after a usage limit resets - #63

Merged
lukemaj merged 4 commits into
mainfrom
feat/61-resume-after-usage-limit
Sep 28, 2026
Merged

lukemaj merged 4 commits into
mainfrom
feat/61-resume-after-usage-limit

Conversation

@lukemaj

@lukemaj lukemaj commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Closes #61.

When a provider's usage limit stopped a thread, nothing brought it back: the user had to return after the reset and type "continue". Three threads stopped this way on 2026-09-28 alone (each turn ended as error about 1 s after starting).

When a thread the user started fails while its provider instance has an exhausted usage window with a known reset (the Usage panel's data), the threads toolkit sends one "Continue. (Sent automatically after the usage limit reset.)" a minute after the latest reset. The decision reads the provider's usage windows only, so it covers every provider that reports them, with no error-text parsing.

  • Starting another turn or archiving the thread before then cancels it.
  • Child threads (sub.*) are skipped: Model Router reroutes or blocks Prism jobs itself, and the planner decides what to do with a blocked job once it resumes.
  • Pending resumes live in memory; a server restart drops them.
  • It hangs off the threads toolkit's existing domain-event watcher, so no upstream file changes. Upstream Orchestrator V2 has its own "Resume at reset" (feat(v2): resume limited threads when usage resets pingdotgg/t3code#12686), so this goes away with the V2 port.

The Elon record (cuts and evidence) is on #61.

Proof

  • vp test run src/mcp/toolkits/threads/usageLimitResume.test.ts: 5 passed on TestClock (resumes once a minute after the reset and not at the reset itself; nothing when not limited; stands down after a newer turn or archive; reset selection across windows).
  • vp test run src/mcp/toolkits/threads/: 36 passed, including the real-engine child report-back test that builds this layer.
  • apps/server typecheck clean; lint clean on changed files; scripts/fork-check.sh OK.
  • Not yet observed on a real limit hit.

Done by Claude Opus 5.5 in Claude Code.

🤖 Generated with Claude Code

lukemaj added a commit that referenced this pull request Sep 28, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Sep 28, 2026
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ The exact PR base did not have a successful artifact. Baseline uses the latest successful main measurement shown below.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire 13.5 KiB 13.5 KiB +21 B (+0.2%) 15.1 KiB ✅
Codex Thread snapshot wire 7.1 KiB 7.1 KiB +1 B (+0.0%) 7.3 KiB ✅
Codex Live turn WebSocket wire 6.4 KiB 6.5 KiB +20 B (+0.3%) 7.8 KiB ✅
Codex Live turn WebSocket decoded 56.2 KiB 56.3 KiB +44 B (+0.1%) 66.4 KiB ✅
Codex Live turn messages 9 10 +1 (+11.1%) 21 ✅
Claude Total thread wire 13.5 KiB 13.5 KiB −13 B (−0.1%) 15.1 KiB ✅
Claude Thread snapshot wire 7.1 KiB 7.1 KiB +1 B (+0.0%) 7.3 KiB ✅
Claude Live turn WebSocket wire 6.4 KiB 6.4 KiB −14 B (−0.2%) 7.8 KiB ✅
Claude Live turn WebSocket decoded 57.0 KiB 57.0 KiB 0 B (0.0%) 66.4 KiB ✅
Claude Live turn messages 9 9 0 (0.0%) 21 ✅

Baseline: 19a04dc · PR result: 25061bd · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 113.9 KiB
  • Claude decoded thread snapshot: 114.6 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

lukemaj and others added 3 commits September 28, 2026 20:11
…61)

When a provider's usage limit stopped a thread, nothing brought it back:
the user had to return after the reset and type "continue". Three
threads stopped this way on 2026-09-28 alone.

When a thread the user started fails while its provider instance has an
exhausted usage window with a known reset (the data the Usage panel
shows), the threads toolkit sends one "Continue. (Sent automatically after
the usage limit reset.)" a minute after the latest reset. Starting another
turn or archiving the thread first cancels it. Child threads are skipped:
Model Router reroutes or blocks Prism jobs itself, and the planner decides
what to do with a blocked job once it resumes. Pending resumes live in
memory, so a server restart drops them. Upstream Orchestrator V2 has its
own "Resume at reset", so this goes away with the V2 port.

Done by Claude Opus 5.5 in Claude Code.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@lukemaj
lukemaj force-pushed the feat/61-resume-after-usage-limit branch from 42e8e1f to 6ffec82 Compare September 28, 2026 18:12
Knip flagged three exports only used inside mobileShell.ts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@lukemaj
lukemaj merged commit ca3706f into main Sep 28, 2026
18 checks passed
@lukemaj
lukemaj deleted the feat/61-resume-after-usage-limit branch September 28, 2026 18:23
lukemaj added a commit that referenced this pull request Oct 1, 2026
…limit (#71)

A spawn_thread child that stopped on a usage limit waited until the user
typed "continue" to the planner and the planner relayed it: #63 skipped
every sub.* thread to leave Prism jobs to Model Router.

The resume watcher now skips only Prism job threads, recognized by Model
Router's "[model-router job ...]" first user message (titles get
regenerated, so they are no marker). A resume fires at the earlier of the
displayed reset plus a minute, or the first reply any thread gets from the
same provider instance after the settle delay: on 2026-09-29 the displayed
reset was 2.5 hours later than when requests worked again. A thread whose
early resume fails again waits for the reset, so a per-model limit cannot
turn replies elsewhere into a retry loop. A resumed child's parent gets
one line saying so.

Done by Claude Opus 5.5 in Claude Code.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
lukemaj added a commit that referenced this pull request Oct 1, 2026
…limit (#71)

A spawn_thread child that stopped on a usage limit waited until the user
typed "continue" to the planner and the planner relayed it: #63 skipped
every sub.* thread to leave Prism jobs to Model Router.

The resume watcher now skips only Prism job threads, recognized by Model
Router's "[model-router job ...]" first user message (titles get
regenerated, so they are no marker). A resume fires at the earlier of the
displayed reset plus a minute, or the first reply any thread gets from the
same provider instance after the settle delay: on 2026-09-29 the displayed
reset was 2.5 hours later than when requests worked again. A thread whose
early resume fails again waits for the reset, so a per-model limit cannot
turn replies elsewhere into a retry loop. A resumed child's parent gets
one line saying so.

Done by Claude Opus 5.5 in Claude Code.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
lukemaj added a commit that referenced this pull request Oct 1, 2026
…limit (#72)

* feat(server): direct child threads resume on their own after a usage limit (#71)

A spawn_thread child that stopped on a usage limit waited until the user
typed "continue" to the planner and the planner relayed it: #63 skipped
every sub.* thread to leave Prism jobs to Model Router.

The resume watcher now skips only Prism job threads, recognized by Model
Router's "[model-router job ...]" first user message (titles get
regenerated, so they are no marker). A resume fires at the earlier of the
displayed reset plus a minute, or the first reply any thread gets from the
same provider instance after the settle delay: on 2026-09-29 the displayed
reset was 2.5 hours later than when requests worked again. A thread whose
early resume fails again waits for the reset, so a per-model limit cannot
turn replies elsewhere into a retry loop. A resumed child's parent gets
one line saying so.

Done by Claude Opus 5.5 in Claude Code.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(server): interrupt_thread cancels a usage-limit auto-continue (#71)

A dispatcher interrupts a limit-hit child before replacing it, but that
child is already failed, so interrupt_thread had nothing to stop while the
pending automatic continue would still restart it next to its replacement.
interrupt_thread now cancels that pending continue and marks the stopped
turn so a later limit error for it schedules none.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(server): continue only usage-limit errors, and past resets keep watching (#71)

Only a turn whose own error reads as a usage limit schedules a continue;
an unrelated runtime error while a usage window is exhausted is left
alone. An exhausted window whose reset time already passed no longer ends
the watch: it keeps waiting for a reply on the instance and re-reads the
windows every 5 minutes, resuming at a future reset or once none is
exhausted, so a stale reading cannot turn into a continue loop.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(server): resume at a reset only once the window shows it lifted (#71)

At a future reset plus the margin the watcher resumed without reading the
usage windows again, so a window still at 100 % got a continue anyway. It
now re-reads them and resumes only when none is exhausted; otherwise it
keeps re-reading every 5 minutes, and a reply on the instance can still
resume it early.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Resume threads after a usage limit resets

1 participant