Skip to content

[Bug]: Grok ACP resends full cumulative terminal output in every tool_call_update, flooding ingestion for all threads #6556

Description

@kelchm

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Start a Grok provider thread and have it run a long shell command that redraws a progress bar (observed with a ~21 GB model download; any high-churn terminal output should work).
  2. Let the command run for several minutes, until its accumulated output exceeds ~100 KB.
  3. Keep one or more unrelated threads, using any provider, active at the same time.
  4. Watch ~/.t3/userdata/logs/provider/events.<grok-thread-id>.log and the responsiveness of the other threads.

Expected behavior

tool_call_update handling should be bounded. The adapter should coalesce or truncate high-frequency terminal updates, or diff them against the previous content, so that a redrawing progress bar produces at most kilobytes of runtime events per second. One busy Grok tool call should not degrade ingestion for unrelated threads.

Actual behavior

The Grok CLI emits a session/update → tool_call_update notification for every progress-bar redraw. Each notification contains the entire accumulated terminal output of the tool call, not a delta.

At a ~10/sec redraw rate with ~145 KB of accumulated output, this peaked at a sustained ~1.1 MB/s of JSON-RPC traffic, with elevated event volume from ~03:22 to ~03:56 UTC on 2026-08-13 (orchestration receipts show the flood driving ~600 ingestion dispatches/min from 03:23):

  • Individual notifications reached 145,615 bytes. Stdio delivered them in 64 KB chunks, and many decode passes completed zero messages, causing even more raw/decoded protocol log churn.
  • At peak the per-thread provider NDJSON log rotated a full 10 MB every ~9 seconds: nine consecutive retained rotations (~90 MB) span just 79 seconds (03:27:23–03:28:42). Earlier rotations from the storm had already been deleted; traffic then slowed (the next file spans 03:28:42–03:55:25).
  • The _meta.eventId suffix reached -22430 during the session, indicating tens of thousands of notifications for one turn.

GrokAdapter canonicalizes these updates without size or rate limits. Because ProviderRuntimeIngestion and the orchestration engine process runtime events serially (see #5681), the flood creates head-of-line blocking: provider events for unrelated threads stop being projected.

In the incident that exposed this problem, ingestion stopped dispatching entirely at 03:26:17 while the storm continued. This froze live updates for every other thread and set up the permanent state loss described in #6560.

This is arguably also an upstream Grok CLI defect because tool_call_update contains cumulative content rather than a delta. The server still needs to defend against it: one provider should not be able to wedge ingestion for the entire app.

Impact

Major degradation or frequent failure

Version or commit

Desktop 0.0.33 / main @ e5c82d7

Environment

macOS 26.5.2 (Darwin 25.5.0), Grok CLI 1.0.3 (grok-4.5)

Logs or stack traces

# Raw stdio frames: 64 KB chunks, ~half of decode passes yield 0 complete messages
[2026-08-13T03:27:23.104Z] NTIVE: {..."kind":"protocol","provider":"grok","payload":{"direction":"incoming","stage":"raw","payload":{"valueType":"string","byteLength":65536}}}
[2026-08-13T03:27:23.104Z] NTIVE: {..."kind":"protocol","provider":"grok","payload":{"direction":"incoming","stage":"decoded","payload":{"valueType":"array","itemCount":0}}}

# The payload: full cumulative progress-bar output in a tool_call_update
[2026-08-13T03:27:23.108Z] NTIVE: {..."method":"session/update",...,"update":{"sessionUpdate":"tool_call_update","content":[{"type":"content","content":{"type":"text","text":"███████       ▏  15 GB/ 21 GB   61%..."}}]}}

# One retained rotation (10 MB / ~8 s of wall clock):
#   415 log lines, 166 raw frames (116 of them ~64 KB), 83 session/update notifications,
#   largest single line 145,615 bytes
# Rotation cadence during the storm:
#   log.10  03:27:23 -> 03:27:31    log.9  03:27:31 -> 03:27:40    log.8  03:27:40 -> 03:27:49 ...

Screenshots, recordings, or supporting files

No response

Workaround

Interrupt the Grok turn, or avoid progress-bar-heavy commands in Grok threads. Once the storm stops, ingestion catches up—unless the server restarts first, which triggers #6560.

Activity

  1. added
    bugSomething is broken or behaving incorrectly.
    needs-triageIssue needs maintainer review and initial categorization.
    on Aug 14, 2026
  2. maslinedwin commented on Aug 16, 2026

    @maslinedwin
    Contributor

    Fix is in #7070 (b0079ede). Grok tool_call_update handling now rate-limits in-progress execute ticks (250ms), skips identical snapshots, and truncates cumulative terminal content to 8KB so a redrawing progress bar cannot flood ingestion for other threads.

  3. VincentShipsIt commented on Aug 22, 2026

    @VincentShipsIt

    Confirming this is still live on T3 Code desktop 0.0.33 / Grok 4.6 (ACP) as of 2026-08-22.

    Thread: Grok /deploy on a Vitae workspace
    T3 thread 1b0fc04c-b87a-4fb2-92b0-c36c386b6128
    Grok session 01a029af-71dc-7930-a00b-f60c783c3eaa

    What we saw

    • T3 projection and events.<thread>.log stop at 15:23:32Z on the exact signature from this issue: incoming raw byteLength: 65536 then decoded itemCount: 0.
    • Grok itself kept running. updates.jsonl continued past that (eventId suffix reached -24630). The ACP process was still alive after T3 Stop (PID still holding the session events.jsonl + a monitor log).
    • UI stayed on Working with user pings (status?) queued as a pending turn that never started. Same shape as Grok turn stalls mid-reasoning: no tokens for 70+ minutes, UI stays Working, no turn_ended #7210.

    Grok native chat_history.jsonl for that session: 167 assistant messages, 147 with empty content (tool_calls only), 20 with real text. T3 only projected the 20 textful ones, so the thread looks blank between user messages even before the wedge. Tool rows show as +N previous tool calls.

    Settled OpenCode threads from the last 7 days have assistant text in state.sqlite (zero empty assistant rows) but look the same in the UI because of #7518.

    Happy to attach sanitized log timestamps if useful. We stopped the T3 thread; the Grok ACP child outlived Stop.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.needs-triageIssue needs maintainer review and initial categorization.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions