What happened
On every launch of a desktop build from the Orchestrator V2 branch (#2829), the Pi provider shows "Pi CLI is installed but timed out while running pi --version." The status clears on the next provider health refresh 5 minutes later. pi --version takes 0.15–0.22 s on the same machine.
The Pi probe is not the problem. A synchronous startup SQLite scan blocks the server event loop for longer than the probe's 4 s timer. The scan's size grows without bound because the V1 projector cursors stop advancing once V2 traffic takes over.
Diagnosis
checkPiProviderStatus starts at 12:56:32.548 with VERSION_PROBE_TIMEOUT_MS = 4_000 (PiProvider.ts).
- At 12:56:32.627,
ProjectionPipeline.bootstrap runs listAttachmentCleanupReplayRows. The span is sql.execute and it takes 5236 ms. The client is node:sqlite DatabaseSync, so the query runs synchronously on the main thread.
pi --version exits after about 150 ms, but the exit is not processed until the query returns. libuv runs expired timers before poll-phase I/O, so the 4 s timer fires first. The probe span ends at 12:56:37.868, 5 ms after the SQL returns, and takes the Option.none "timed out" branch.
- The next refresh, at 13:01:32 (
DEFAULT_PROVIDER_HEALTH_REFRESH_INTERVAL = 5 min), succeeds.
Why the scan is large. The bootstrap scans from cleanupStart = min(cleanup cursor, every V1 projector cursor). The V1 projectors read through readFromSequence, which returns only application_event_version = 1 OR aggregate_kind = 'project'. Their cursors move only when a project event arrives, and V2 thread traffic never moves them.
projection_state on one machine:
projection.projects … projection.threads 58658 (2026-09-21, the last project event)
projection.attachment-cleanup 427880
max(sequence) 428271
So cleanupStart is 58658, and every restart re-reads about 370k V2 rows with the NOT INDEXED rowid scan. That window grows with all traffic since the last project change, not with traffic since the previous restart, which is what #12846 aimed for. Timing the same query read-only with the sqlite3 CLI gives 2.3 s cold / 0.36 s warm from 58658, and 1 ms from a current cursor.
A second machine has the same pinning (cursors at 277538, head 714625), but it is fast enough (M5 Max, about 0.3 s) to stay under 4 s. The affected machine is an M2 Pro with a 2.7 GB statev2.sqlite and endpoint security (CrowdStrike, Defender), which pushes the cold scan past the 4 s timer.
Steps to reproduce
- Use a V2 build with some thread history. Create no project events afterwards, so the V1 projector cursors stay behind the head.
- Let V2 thread traffic accumulate.
- Relaunch. When the cleanup scan exceeds 4 s on the machine, the Pi provider shows the timeout error for one refresh interval.
Mechanism in isolation (Node, no T3): spawn pi --version with a 4 s setTimeout, then busy-block the loop.
block 0ms → ok 0.87.1 at 109ms
block 3500ms → ok 0.87.1 at 3502ms
block 5200ms → TIMED OUT at 5202ms
Version
Built from #2829 head 402205e2c7 (0.0.42, Electron 44.4.2), Pi 0.87.1.
Environment
macOS 26.6.2, Apple M2 Pro, 32 GB. EDR installed: CrowdStrike Falcon, Microsoft Defender, Kandji ESF.
Evidence
server.trace.ndjson spans from the launch at 12:56:25:
12:56:32.548 5320 ms checkPiProviderStatus (ends 5 ms after the SQL returns; no discovery phase)
12:56:32.627 5236 ms sql.execute SELECT sequence, occurred_at … FROM orchestration_events NOT INDEXED WHERE sequence > 58658 …
13:01:32.551 checkPiProviderStatus (next refresh, succeeds)
Related issues
#12846 (perf: stop decoding unrelated events during startup) added the separate cleanup cursor. The min() over projector cursors that stop advancing under V2 undoes that bound.
Fix applied or workaround
A local patch (with a regression test) is applied in a fork: astarktc@587ae54
At the end of bootstrap's projector replay, it advances every projector cursor still below the pre-replay log head up to that head. The replay has consumed everything through the head; the V2 events were simply filtered out. The next cleanup scan window is then bounded by the traffic since the previous start. This needs no migration.
Suggested follow-ups, in order of value:
- Make the cleanup query independent of history size. A partial index such as
ON orchestration_events(sequence) WHERE aggregate_kind = 'thread' AND event_type IN ('thread.reverted','thread.deleted'), plus MAX(sequence) from the rowid, would make the scan O(reverts + deletes) regardless of cursors. That needs a migration, which is why the fork patch does not do it.
- Treat a timed-out health probe as retryable. Re-check after a short backoff instead of holding a false error for the full 5-minute interval. Any long synchronous stall at startup, not only this query, currently produces this false error.
Filed by
@astarktc, with an AI agent (Pi in T3 Code), after a trace-level diagnosis on the affected machine.
What happened
On every launch of a desktop build from the Orchestrator V2 branch (#2829), the Pi provider shows "Pi CLI is installed but timed out while running
pi --version." The status clears on the next provider health refresh 5 minutes later.pi --versiontakes 0.15–0.22 s on the same machine.The Pi probe is not the problem. A synchronous startup SQLite scan blocks the server event loop for longer than the probe's 4 s timer. The scan's size grows without bound because the V1 projector cursors stop advancing once V2 traffic takes over.
Diagnosis
checkPiProviderStatusstarts at 12:56:32.548 withVERSION_PROBE_TIMEOUT_MS = 4_000(PiProvider.ts).ProjectionPipeline.bootstraprunslistAttachmentCleanupReplayRows. The span issql.executeand it takes 5236 ms. The client isnode:sqliteDatabaseSync, so the query runs synchronously on the main thread.pi --versionexits after about 150 ms, but the exit is not processed until the query returns. libuv runs expired timers before poll-phase I/O, so the 4 s timer fires first. The probe span ends at 12:56:37.868, 5 ms after the SQL returns, and takes theOption.none"timed out" branch.DEFAULT_PROVIDER_HEALTH_REFRESH_INTERVAL= 5 min), succeeds.Why the scan is large. The bootstrap scans from
cleanupStart = min(cleanup cursor, every V1 projector cursor). The V1 projectors read throughreadFromSequence, which returns onlyapplication_event_version = 1 OR aggregate_kind = 'project'. Their cursors move only when a project event arrives, and V2 thread traffic never moves them.projection_stateon one machine:So
cleanupStartis 58658, and every restart re-reads about 370k V2 rows with theNOT INDEXEDrowid scan. That window grows with all traffic since the last project change, not with traffic since the previous restart, which is what #12846 aimed for. Timing the same query read-only with the sqlite3 CLI gives 2.3 s cold / 0.36 s warm from 58658, and 1 ms from a current cursor.A second machine has the same pinning (cursors at 277538, head 714625), but it is fast enough (M5 Max, about 0.3 s) to stay under 4 s. The affected machine is an M2 Pro with a 2.7 GB
statev2.sqliteand endpoint security (CrowdStrike, Defender), which pushes the cold scan past the 4 s timer.Steps to reproduce
Mechanism in isolation (Node, no T3): spawn
pi --versionwith a 4 ssetTimeout, then busy-block the loop.Version
Built from #2829 head
402205e2c7(0.0.42, Electron 44.4.2), Pi 0.87.1.Environment
macOS 26.6.2, Apple M2 Pro, 32 GB. EDR installed: CrowdStrike Falcon, Microsoft Defender, Kandji ESF.
Evidence
server.trace.ndjsonspans from the launch at 12:56:25:Related issues
#12846 (perf: stop decoding unrelated events during startup) added the separate cleanup cursor. The
min()over projector cursors that stop advancing under V2 undoes that bound.Fix applied or workaround
A local patch (with a regression test) is applied in a fork: astarktc@587ae54
At the end of
bootstrap's projector replay, it advances every projector cursor still below the pre-replay log head up to that head. The replay has consumed everything through the head; the V2 events were simply filtered out. The next cleanup scan window is then bounded by the traffic since the previous start. This needs no migration.Suggested follow-ups, in order of value:
ON orchestration_events(sequence) WHERE aggregate_kind = 'thread' AND event_type IN ('thread.reverted','thread.deleted'), plusMAX(sequence)from the rowid, would make the scan O(reverts + deletes) regardless of cursors. That needs a migration, which is why the fork patch does not do it.Filed by
@astarktc, with an AI agent (Pi in T3 Code), after a trace-level diagnosis on the affected machine.