Skip to content

[Bug]: Claude context meter jumps to 100% at end of turn - completeTurn falls back to cumulative session usage from result.usage #8594

Description

@cristip73

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Run a Claude thread (1M context window) and let a turn do real work, so it makes several tool calls.
  2. Let the turn finish normally, without interrupting it.
  3. Watch the context meter in the composer at the moment the turn ends.

The turn has to complete on its own. Interrupted turns take a different branch in completeTurn and are not affected.

This reproduces reliably when query.getContextUsage() is unavailable or slow. On my machines that call takes 780-1152 ms (n=51, median 880) on a Mac mini and 1069-1542 ms (n=7, all over budget) on a MacBook, against the 1 second budget at ClaudeAdapter.ts:2150. On the MacBook it therefore misses on essentially every turn.

Expected behavior

The meter keeps showing the active context, roughly the value it displayed during the turn.

Actual behavior

At the end of the turn the meter jumps to exactly maxTokens and shows 100% - 1m/1m, together with the "Resume with less context" banner, on a thread whose real context is around 113k. It stays wrong until a later turn happens to get an authoritative reading.

The chain, all in apps/server/src/provider/Layers/ClaudeAdapter.ts:

  1. completeTurn asks the authoritative source first: queryCurrentContextUsage -> context.query.getContextUsage(), wrapped in Effect.timeoutOption("1 second") (~:2150). Failure is silent: the catch at ~:2147 is empty and timeoutOption swallows the timeout, so nothing is logged.
  2. On failure it falls back to result.usage through normalizeClaudeActiveTokenUsage (~:2277).
  3. result.usage is not the active context. The CLI builds it by summing the per-model accumulators in modelUsage, which are never reset, so it accumulates across the whole session. claudeUsageInputTokens (~:505) then adds input + cache_creation + cache_read.
  4. makeClaudeTokenUsageSnapshot clamps with usedTokens = Math.min(activeTokens, maxTokens) (~:557), producing exactly maxTokens.

Two stored context-window.updated payloads from the same thread, three seconds apart:

{"usedTokens":112994,"inputTokens":111443,"maxTokens":1000000}
{"usedTokens":1000000,"totalProcessedTokens":2202960,"inputTokens":2184473,"outputTokens":18487,"maxTokens":1000000}

2184473 + 18487 = 2202960, which is exactly the cumulative total. The number rendered as "context used" is that total, clamped.

In one week on a single install, 93 stored events carry usedTokens of exactly 1,000,000. The most inflated had a cumulative inputTokens of 19,767,501.

A note on the iterations guard

resultHasActiveUsage (~:2269) is widened by hasResultUsageIteration so that per-iteration snapshots keep updating the meter. The bundled CLI never emits them: its result usage template is {..., iterations: [], speed: "standard"} and the aggregator only fills the token counters, so lastClaudeUsageIteration always returns undefined for result.usage. The guard protects a shape that does not occur, while the cumulative path it lets through is the one that breaks the meter.

Relationship to existing issues

This is a different write path from #4650, which is the task_progress ratchet through Math.max. This one is the end-of-turn fallback in completeTurn, and it happens on turns with no task_progress involvement at all. #5942 was closed as a duplicate of #4650 and describes the subagent case, which is also distinct.

#6586 targets the same meter from the subagent angle and has been inactive since 2026-08-14 with a conflicting branch.

Impact

Major degradation or frequent failure

Version or commit

0.0.36-nightly.20260828.1210, source read at main @ ac3b2ad

Environment

macOS 15 (Apple Silicon). Provider: Claude Agent SDK 0.3.221 with the bundled Claude Code 2.1.221. Models with a 1M context window. Reproduced on both a headless t3 serve on a Mac mini and the desktop app's embedded server on a MacBook.

Logs or stack traces

# Nothing is logged when the authoritative read fails. That is the point:
# apps/server/src/provider/Layers/ClaudeAdapter.ts:2143-2150
#   const usage = yield* Effect.promise(async () => {
#     try { return await context.query.getContextUsage?.(); }
#     catch { return undefined; }
#   }).pipe(Effect.timeoutOption("1 second"));
#
# The empty catch also swallows the remote case; the CLI carries the string
# "Context usage isn't available over this remote connection".

Workaround

None from the UI. Patching the bundle locally to raise the budget restores correct readings (7 of 7 turns on a live server emitted the authoritative snapshot), which confirms the mechanism, but raising that budget is not a good fix on its own: that timeout also sits on the Stop path via interruptTurn -> stopSessionInternal -> completeTurn, so a larger value delays teardown after the user presses Stop.

A better direction, if useful: prefer the last per-request reading over the cumulative one. With includePartialMessages: true, every parent message_delta already updates context.lastKnownTokenUsage through normalizeClaudeActiveTokenUsage (~:2473), and BetaMessageDeltaUsage is cumulative per API request, so input_tokens + cache_read_input_tokens there is the real context at the last request of the turn. Keeping result.usage for totalProcessedTokens only would leave that field correct while the meter stops lying.

Happy to open a PR along those lines if it would be welcome.

Activity

  1. cristip73 commented on Aug 29, 2026

    @cristip73
    Author

    Status update after #8610 landed, with measurements from two independent installs.

    The path this report describes is gone from main

    #8610 (c131f289, merged 2026-08-29 03:27Z) removed the query.getContextUsage() call from completeTurn entirely, for an unrelated reason: its token-count fallback could trigger extra model requests. So the 1 second budget I measured no longer exists, and neither does the intermittency that came with it. What took over as the primary source is latestAssistantSnapshot, the last main-thread assistant message's usage, which is per-request and therefore the correct active context. That is the same direction I suggested at the end of this report, reached independently and for a different reason.

    I can no longer reproduce the symptom on nightly

    Counted in orchestration_events, over context-window.updated activities where usedTokens == maxTokens, across two installs (desktop app on a MacBook, headless t3 serve on a Mac mini):

    window events saturated note
    through 2026-08-23, back to February 23,227 0
    2026-08-24 to the upgrade 3,327 118 (3.5%)
    after upgrading to 0.0.37-nightly.20260829.1219 298 0 peak usedTokens 296,989

    The post-fix window is short: about 90 minutes of real use on one machine and 30 on the other. At the prior rate it would predict roughly 10 saturated events, and there are none. I would call that consistent with the fix rather than proof of it, but the symptom that prompted this report is not showing up any more.

    The regression window is narrow, and it is not where I assumed

    The first saturated event of this shape is 2026-08-24T08:23:14Z on one machine and 08:30:58Z on the other, two installs, eight minutes apart, with nothing in the preceding six months of stored events. So this was introduced within a day or so of 2026-08-24 rather than being long-standing.

    A detail that may help anyone testing this

    The two known meter bugs are distinguishable in the stored payload, which I did not realise when filing:

    Splitting on that, my own data separates cleanly: all 118 events above are the completeTurn shape, while the only task_progress saturations I have are 36 events on 2026-08-02, on one machine, none since. Same visible symptom, two different write paths. Worth keeping in mind when writing a regression test for either one, since a test that seeds usage through task_progress is exercising the other path.

    What is still open

    resultIterationSnapshot, built from cumulative result.usage, is still second in the precedence chain in completeTurn. Reading main, it is reachable when latestAssistantSnapshot is undefined and no compaction happened in the turn, which is a much narrower window than before but the same clamp to maxTokens when it fires. I have not observed it in the data above, so I cannot say how often it happens in practice. #8617 targets exactly that residual.

    Happy for this to be closed if you consider the residual covered elsewhere. Thanks for the quick turnaround on #8610.

  2. Mina-Sayed commented on Aug 29, 2026

    @Mina-Sayed

    Thank you for the precise follow-up and the two-install measurements — very helpful.

    You're right that #8610 (c131f28) removed query.getContextUsage() and made latestAssistantSnapshot the primary source, which is the same direction you suggested and for an independent reason (avoiding extra model requests). That explains why the symptom is no longer reproducible on 0.0.37-nightly.20260829.1219 in your data.

    As you note, the residual is still reachable in completeTurn on main: when latestAssistantSnapshot is undefined and no compaction happened in the turn, resultIterationSnapshot (built from cumulative result.usage) is still second in the chain and would clamp to maxTokens. It's a much narrower window than before, but the same incorrect clamp when it fires.

    #8617 targets exactly that residual: latestAssistantSnapshot ?? updatedLastGood ?? (compacted ? undefined : resultIterationSnapshot) — i.e. keep result.usage for totalProcessedTokens/maxTokens only, and preserve the per-request active context from lastKnownTokenUsage (seeded via message_delta / task_progress). Added a focused regression test seeding lastKnownTokenUsage=112994 then delivering a cumulative result 2_202_960 with maxTokens=1M to pin the precedence.

    Thanks again for the quick turnaround on the analysis and for distinguishing the two payload shapes (inputTokens/outputTokens vs toolUses/durationMs) — that detail is now reflected in the test.
    — via #8617 (Mina-Sayed:fix/claude-meter-8594)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions