Skip to content

[Bug]: Streaming assistant deltas rescan full thread activity history #4008

Description

@Chrrxs

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Open a long-running thread with substantial tool/activity history.
  2. Start a turn that streams assistant text.
  3. Observe text delivery becoming dramatically slower than a fresh thread.

Deterministic repro: project a streaming assistant thread.message-sent event against a thread whose activity history contains a decode sentinel. The event currently touches the sentinel through ProjectionThreadActivityRepository.listByThreadId, proving that every delta reads the full activity history.

Expected behavior

Streaming assistant deltas should persist the message and update the thread timestamp without recomputing shell-summary fields that cannot change from assistant text.

Actual behavior

Every streamed assistant delta calls refreshThreadShellSummary, which loads all messages, proposed plans, activities (including full tool payloads), and pending approvals.

On the affected thread, 25,834 activities contained about 254 MB of payloads. The provider emitted 55 deltas in 521 ms, while server-side assistant-delta projection took a median 819 ms per delta (p95 1,041 ms), making text visibly trickle.

This is distinct from #2761 (snapshot/WebSocket congestion) and #4005 (client cache writes).

Impact

Major degradation or frequent failure

Version or commit

main at ecb35f758

Environment

Windows 11 with WSL2, T3 Code desktop/nightly, Node.js 24.16.0, Codex provider.

Logs or stack traces

activities: 25,834
activity payloads: ~254 MB
assistant.delta projection: median 819 ms, p95 1,041 ms
provider output: 55 deltas / 251 chars in 521 ms

Screenshots, recordings, or supporting files

T3Code.Slow.TPS.Issue.2026-07-15.100931.mp4

Workaround

Start a fresh thread. A local fix that skips the shell-summary refresh only for streaming assistant messages restored normal text delivery while preserving message projection and updatedAt.

Activity

  1. VincentShipsIt commented on Jul 30, 2026

    @VincentShipsIt

    Desktop repro data (macOS, packaged app) — matches this issue closely

    Environment

    • T3 Code (Alpha) 0.0.31 (com.t3tools.t3code)
    • macOS desktop, local backend on 127.0.0.1:3773
    • Grok Build (ACP) + Codex + Claude Agent all configured; slowdown reproduced during multi-tool Grok turns
    • enableAssistantStreaming: true in ~/.t3/userdata/settings.json at capture time

    Local profile at capture (~/.t3/userdata/state.sqlite)

    Metric Value
    DB file size 587 MB
    orchestration_events 151,749
    thread.message-sent 127,380 (~84%) / ~38 MB payload
    thread.activity-appended 22,813 / ~174 MB payload
    Projected activities 22,813 / ~168 MB
    Projected messages 2,286 / ~0.5 MB

    Heaviest threads (pre-trim)

    Title Activities Events Status
    tests coverage 14,709 102,002 archived+deleted (still on disk)
    QA 2,461 (~80 MB activities) 11,159 active
    /deploy 2,246 16,462 archived+deleted
    Fix Messages Link Routing 1,062 4,713 active
    QA (other) 981 (~20 MB) 5,688 active

    User-visible symptom
    Grok Build turns start fine, then after a couple of tool-call rounds in the same thread the UI becomes dramatically slower (text trickle / long gaps between tool steps). Fresh threads feel normal; long agentic threads degrade mid-turn. This is not “Grok CLI is slow” — terminal Grok is fine; T3 falls over as history grows.

    Why this looks like #4008

    • Streaming was on → event log dominated by thread.message-sent deltas (one event per stream chunk).
    • Tool rounds add large thread.activity-appended payloads (activity payload mass ≫ message text mass).
    • Slowdown only appears after the thread accumulates activities, consistent with per-delta work that re-touches full activity history / shell summary (as described here), not with provider token latency.

    Server traces (supporting)

    • Large volumes of runProjectorForEvent + sql.execute during agent turns
    • Long-lived ws.rpc.orchestration.subscribeThread spans (minutes–hours class in duration fields)
    • Separate Grok ACP noise also present historically (Cannot convert skills-reload to a BigInt / model discovery timeout) — orthogonal, not the mid-turn regression

    Workaround that recovered the machine

    1. Full quit of the app
    2. SQLite .backup, then purge soft-deleted thread rows + wipe history on fat active threads (kept thread shells + provider session rows)
    3. Set enableAssistantStreaming: false
    4. VACUUM → DB 587 MB → 288 KB

    Happy to re-capture timing of assistant-delta projection (median/p95) on a synthetic thread if useful for a regression test. Cross-links: also matches client quadratic replay in #4596 and unbounded history hydration in #2761 / #3510.

  2. DTavaszi commented on Aug 4, 2026

    @DTavaszi

    Confirmed a broader variant of this issue with assistant streaming disabled.

    A long-lived thread containing tens of thousands of activities made each thread.activity-appended dispatch spend roughly two seconds in refreshThreadShellSummary. ProviderRuntimeIngestion uses a single serial DrainableWorker for all provider events, so these slow events created cross-thread head-of-line blocking: an unrelated thread completed normally at the provider, but its response was not persisted and rendered by T3 Code for about seven minutes.

    This suggests the impact is not limited to assistant-delta streaming within the large thread. Routine activity events from one large thread can starve unrelated conversations globally.

    The completion was not lost or corrupted; it became visible after the ingestion backlog drained. In addition to making shell-summary projection incremental, partitioning ingestion by thread or moving expensive projections off the global serial path may prevent the cross-thread failure mode.

  3. ognjeeen commented on Aug 6, 2026

    @ognjeeen

    Still reproducible on T3 Code 0.0.31 across two independent Windows installations.

    Additional measurements and controlled cleanup A/B data are documented in #5110:

    • one 10,545-character response produced 3,024 durable thread.message-sent events (median delta: 3 characters);
    • after cleanup, one database regrew from 282 KB to 218 MB and 93,943 events in approximately 25 hours;
    • removing only T3 Code's internal history immediately restored responsive streaming;
    • PRAGMA quick_check remained ok;
    • built-in thread deletion did not purge historical events, receipts, messages, or activities.

    This supports #4008 while suggesting that the shell-summary rescan is one amplifier within a broader per-delta persistence, queueing, and retention problem.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.🚧 In Progress

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions