Repository navigation
[Bug]: Bounded thread snapshot returns the whole thread when its runs start from notifications (PR watch, task wakes) #17433
Description
Activity
Note
Grok responding on behalf of Julius.
Thanks for the detailed numbers. I haven't run this against a large database; below is what current
main(43f8a8de1) does on the paths you named, and it lines up with your reading.Notification-started runs don't count as turn starts
notificationTurnItemrewrites auser_messagethat carries anotification/delegatedCompletionintotype: "notification", and stripsinputIntent,createdBy,messageIdandtextin the process (L224-L234). The Orchestrator persists that rewritten payload at four call sites (Orchestrator.ts#L1839, L3971, L5506, L6409).isThreadHistoryTurnStartonly matchesuser_messagewithinputIntentturn_start/queued_turn, so a run whose root item is anotificationisn't an anchor for either cap.
SQL window (
readCanonicalProjection)turn_anchorstakes up toTHREAD_HISTORY_MAX_RAW_TURNS + 2(152, with the cap at 150) user-message turn starts;user_anchorskeeps the newestuserTurnLimit + 2user-created ones (12 for the first load, sincemaxUserTurnsis 10).boundaryuses the oldest of those user anchors when there are enough of them, else the oldest of 152 raw anchors, else0.selectedthen usesLIMIT -1as soon asuser_anchorsis non-empty (L2815). TherowLimit(77) only applies to histories with no user turn starts.
That looks deliberate for ordinary threads: the policy comment says item/byte budgets only apply to histories without turn starts, so tool activity doesn't split a turn, and the raw-turn cap is meant to keep agent-driven turns from stretching the window. Both rely on every run starting with an anchor, though. In your thread the 3,759 PR-watch wakes are invisible to
turn_anchors, so the raw cap seems unlikely to trip, the boundary lands on the 12th-newest user turn near the start, andLIMIT -1returns everything after it. Nothing in this query bounds item count or bytes once a user anchor exists.In-memory path
layerMemory.getThreadSnapshotWindowmirrors the SQL: it counts anchors withisThreadHistoryTurnStart, and when anchors exist it slices from a turn boundary (orrawStart, which is0below 152 anchors) with no row cap.selectOlderTimelinePagealso drops the item/byte budget as soon as any user turn exists (L212) and only incrementsrawTurnsonisThreadHistoryTurnStart, so older pages appear to have the same gap.Recent changes
#17029 (merged) added these window CTEs and the
user_messagepartial index in060_ThreadSnapshotWindowIndexes.ts. #17387 (merged) changed parts of this query andbuildBoundedThreadProjection's budget check, but the anchor predicate and theLIMIT -1branch are unchanged onmain. Open PR #16819 narrowsLIMIT -1to a positive boundary (the single-open-turn case); as far as I can tell it keeps the positive-boundary branch unbounded, so it likely wouldn't cover this thread on its own.Possible fixes (not mutually exclusive)
- Count notification-rooted runs as turn starts. Treat a run-root
notificationitem (or preserveinputIntenton it) as a raw-turn anchor inisThreadHistoryTurnStartand in theturn_anchorsSQL. The 152 raw-turn cap would then likely bound your thread. Caveats: the partial index only coverstype = 'user_message', so the CTE would need an index that includes notifications; and it's worth checking that everynotificationitem is a run root before relying on the type alone (your DB says yes, I haven't checked other delivery paths). - Add an item/byte ceiling to bounded windows. Stop using
LIMIT -1for a bounded snapshot: cap rows (or bytes) even when user anchors exist, and hand backhasMoreHistory/ a cursor for the rest. That would also guard against other anchor-poor histories, at the cost of sometimes splitting a turn, which the current policy tries to avoid.
Related
- #14701 — thread list read blocking the loop (different query).
- #15365 — iOS reconnect loop when opening one thread; possibly the same symptom on a different cause.
- #16819 — open, same
LIMIT -1branch, single-open-turn case. - #17029, #17387 — merged long-thread snapshot work.
I didn't find a duplicate issue or an open PR that changes how notification-started runs are counted.
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.via-triageFiled through npx t3 triageFiled through npx t3 triage
on Oct 9, 2026 Opened #17441 with fix 1. It handles both caveats:
- Index: wake anchors come from the newest 152 runs through the runs index and each run's root node, so there is no new index or migration.
- Run roots: only a notification that starts its run counts.
appendSteeringMessagecan steer one into a running turn, and that one is excluded.
On a copy of the database, the bounded first load of the reported thread went from 20,035 rows (61 MB, 6.6-8.4 s, server unresponsive about 3 s) to 75 rows (0.26 MB, 0.15-0.21 s). The PR has the full before/after.
Before submitting
Related, but different causes: #14701 (thread list read blocks the loop) and #15365 (reconnects when opening one existing thread). #17387 already fixed the per-item visibility scan on this path. This report is about what is still slow after #17387.
Area
apps/server
Steps to reproduce
watch_pull_requeston the PRs it opens. In four days it got 3,896 runs, of which 3,759 were PR-watch wakeups, plus 15 user turns at the start.GET /api/orchestration/threads/<id>, or reconnect a client that has it open.Expected behavior
The first load returns the bounded window: the latest user turns, capped at
THREAD_HISTORY_MAX_RAW_TURNS(150) turns. Older history pages in on demand.Actual behavior
The "bounded" snapshot is the whole thread: 22,714 turn items, 100 MB of JSON. Building and sending it keeps the server from answering anything else for 3.6 to 4.1 s. A client with the thread open drops on the stall, reconnects, and reloads the thread, so the disconnects repeat for as long as the thread stays open.
Cause, in
readCanonicalProjection(SQLturn_anchors) andisThreadHistoryTurnStart:notificationTurnIteminto anotificationitem, which dropsinputIntent. So a notification-started run is not a turn start.turn_anchorstherefore only sees the 85 user-message turn starts in this thread, below the 152 needed for the raw-turn cap. 15 of them are user turns, which is at leastuserTurnLimit + 2, so the boundary is the 12th-newest user turn, near the start of the thread.user_anchorsis non-empty,selectedhasLIMIT -1. Every row from that boundary on is returned.ProjectionStore.layerMemory'sgetThreadSnapshotWindowandselectOlderTimelinePagecount turns withisThreadHistoryTurnStartand have the same gap.Impact
Major degradation or frequent failure
Version or commit
mainat 43f8a8d, also0.0.46-nightly.20261009.2861Environment
Linux x64 server (
t3 serve, web mode), Node 24,statev2.sqlite4.2 GBLogs or stack traces
The server was run against a copy of the database in a sandbox with no network, and a client loaded the thread four times per path. "Unresponsive" is the longest time a 50 ms HTTP probe waited for any answer.
Item types in that thread: 3,782
notification, 100user_message(12turn_startand 3queued_turnfrom the user, the rest from the agent or steers). In the whole database, all 7,200notificationitems are the root item of their run.Note
Claude (Opus 5.5, Claude Code) wrote this report for @EldarPolitkin.