Before submitting
Related, not duplicates: #16217 / #16208 (subagents woken by inherited watches; fixed, and my subagents had no wakes), #13988 (resume compaction offered only after the cache expired), #10955 (Stop re-bills the cache).
Area
apps/server (PR watch, orchestration-v2/pullRequestWatch.ts)
Steps to reproduce
- In a Claude thread, work until the context is large (mine: Opus 5.5, 1M window, ~760k tokens) and let the agent
watch_pull_request on its PR.
- Leave the thread for more than an hour (the Claude subscription prompt cache lives 1h).
- Someone comments on the PR, or its checks finish.
- Compare
cache_creation_input_tokens / cache_read_input_tokens of the woken turn in ~/.claude/projects/<project>/<session>.jsonl, and the five_hour utilization in userdata/logs/provider/events.*.log.
Expected behavior
A wake does not silently start an expensive turn on a thread whose context is large and whose cache has expired. For example:
- above a size threshold with a cold cache, compact first, or notify the user instead of starting a turn;
- no wake for news that needs no action ("All 1 check passed" when the agent was already done);
- bursts within a few seconds coalesce into one wake.
Actual behavior
The watcher started a turn on the 776k-token thread 18h after its last request. Per-request usage from the transcript:
| local time |
event |
cache_read |
cache_creation |
| 08:25:33 |
PR watch queues "Update on pull request … 7 new comments" |
|
|
| 08:25:47 |
first request |
11k |
750k |
| 08:26–08:29 |
15 more requests |
~762–776k each |
~0 |
| 08:30 |
I interrupt and /compact (thread → ~100k) |
|
|
The provider log's five_hour.utilization went 0.00 → 0.06 on that first request and 0.08 by 08:29, before I had sent anything (Max 5x, new window). The wake turn cost ~2.7M price-weighted input-token units; after my compact the same work ran at ~100–190k context.
The same thread got 6 earlier wakes the day before, each re-reading 557–759k of context (~0.12–0.16M units each): "All 1 check passed" ×4 and "Checks failed" ×2, three of them within 33 seconds (21:01:49, 21:02:06, 21:02:22 UTC).
Summary for this thread: 8 PR wakes = 3.45M units. Native Claude Code has no watcher, so this cost is T3-only.
Impact
Major for long threads that watch a PR: one wake on a large cold thread can take ~8% of a Max 5x 5h window with no user action. Users who compact on return are too late, because the wake fires first.
Version or commit
T3 Code Nightly 0.0.46-nightly.20261005.2702 (desktop, Linux). Bundled Claude Code 2.1.287.
Environment
Linux (Arch/Omarchy), T3 desktop app. Claude subscription (Max 5x, 1h prompt cache). Model Claude Opus 5.5, medium effort, 1M context. PR on a self-hosted Gitea.
Logs or stack traces
# transcript, deduplicated by requestId
2026-10-07T11:25:33Z queue-operation "Update on pull request #NNNN (...), which T3 Code is watching for you: - 7 new comments ..."
2026-10-07T11:25:47Z assistant claude-opus-5-5 cache_read=11k cache_creation=750k
2026-10-07T11:26:02Z assistant claude-opus-5-5 cache_read=762k cache_creation=1k
... 14 more at 763k-776k ...
2026-10-07T11:30:05Z user "[Request interrupted by user for tool use]"
2026-10-07T11:30:11Z /compact
# provider log five_hour utilization (same window)
11:25Z 0.00 -> 11:26Z 0.06 -> 11:27Z 0.07 -> 11:29Z 0.08
Workaround
Compact large threads before stepping away (cheap while the cache is warm), or unwatch the PR for the night.
Before submitting
Related, not duplicates: #16217 / #16208 (subagents woken by inherited watches; fixed, and my subagents had no wakes), #13988 (resume compaction offered only after the cache expired), #10955 (Stop re-bills the cache).
Area
apps/server (PR watch,
orchestration-v2/pullRequestWatch.ts)Steps to reproduce
watch_pull_requeston its PR.cache_creation_input_tokens/cache_read_input_tokensof the woken turn in~/.claude/projects/<project>/<session>.jsonl, and thefive_hourutilization inuserdata/logs/provider/events.*.log.Expected behavior
A wake does not silently start an expensive turn on a thread whose context is large and whose cache has expired. For example:
Actual behavior
The watcher started a turn on the 776k-token thread 18h after its last request. Per-request usage from the transcript:
/compact(thread → ~100k)The provider log's
five_hour.utilizationwent 0.00 → 0.06 on that first request and 0.08 by 08:29, before I had sent anything (Max 5x, new window). The wake turn cost ~2.7M price-weighted input-token units; after my compact the same work ran at ~100–190k context.The same thread got 6 earlier wakes the day before, each re-reading 557–759k of context (~0.12–0.16M units each): "All 1 check passed" ×4 and "Checks failed" ×2, three of them within 33 seconds (21:01:49, 21:02:06, 21:02:22 UTC).
Summary for this thread: 8 PR wakes = 3.45M units. Native Claude Code has no watcher, so this cost is T3-only.
Impact
Major for long threads that watch a PR: one wake on a large cold thread can take ~8% of a Max 5x 5h window with no user action. Users who compact on return are too late, because the wake fires first.
Version or commit
T3 Code Nightly 0.0.46-nightly.20261005.2702 (desktop, Linux). Bundled Claude Code 2.1.287.
Environment
Linux (Arch/Omarchy), T3 desktop app. Claude subscription (Max 5x, 1h prompt cache). Model Claude Opus 5.5, medium effort, 1M context. PR on a self-hosted Gitea.
Logs or stack traces
Workaround
Compact large threads before stepping away (cheap while the cache is warm), or unwatch the PR for the night.