Skip to content

[Bug]: "Resume with less context" is offered only after the prompt cache has expired, so it always takes the expensive path #13988

Description

@Vantrongs

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/web

Steps to reproduce

  1. Work in a Claude thread until it holds more than 100K tokens.
  2. Leave it idle for 70 minutes.
  3. Come back: "Resume with less context" appears.

Expected behavior

The cheap moment is used, and the offer covers Codex too:

  1. Before the 1-hour cache expires, idle threads above a size threshold get a compaction offer: as a notification, since the user is usually away, or as an opt-in automatic step in the thread's live process.
  2. After expiry, the banner also offers "continue in a new thread" with a short brief and a link to the old one; the old thread is archived after the new one's first reply.
  3. The TTL comes from the API response (usage.cache_creation.ephemeral_1h_input_tokens / ephemeral_5m_input_tokens; it is 5 minutes on API keys and in overage). For Codex (gpt-6-astra), 30 minutes is treated as a minimum, and Codex threads get the offer as well (T3 can already compact them with compactThread).
  4. For long agent work, more "Auto-compact after" presets (400K, 600K) and a separate "suggest compaction from N tokens" setting, so a long context stays available when it is needed.

Actual behavior

The banner appears only after 70 idle minutes (Claude Code's own resume threshold, CLAUDE_CODE_RESUME_THRESHOLD_MINUTES), while the cache lives 60 minutes on a Claude subscription and 5 minutes on an API key. By then the cache has expired, so Compact re-reads the whole context at the uncached rate. Codex threads get no offer.

Claude Code's own mechanisms do not cover this. It backs the same 70-minute / 100K dialog with a summary prepared while the cache was warm, and has an idle compaction at 90% of the cache window, but both are server-side experiments: only Anthropic decides which accounts get them, as part of its own testing, and a user has no way to turn them on or influence this. My account (2.1.283) is not enrolled in precomputation (tengu_sepia_moth) or idle compaction (tengu_sunny_locket), but is enrolled in tengu_gleaming_fair_reuse, which shows the dialog only when a precomputed summary exists, so it never appears. Even with the experiment on, the summary is prepared only near the auto-compact threshold.

Measured on Opus 5.5, two idle threads of ~210K tokens, API prices:

Step Cost
Continue a cold thread (the first message writes the context to the 1-hour cache at 2x input) $1.67, then ~$0.05 per message
Compact a cold thread: today's banner, shown after 70 idle minutes while the cache lives 60 $0.90
Compact a warm thread $0.10
First message after either compaction $0.13–0.17, then ~$0.01
New thread with a short brief and a link to the old transcript $0.19 for three answers

Modelled from these numbers, returning to a cold 430K thread for 10 short messages costs $4.40 if continued, $2.12 with today's banner, $0.49 if compacted while warm, and $0.53 via a new thread with a brief and a link. For heavy agent work (~87 requests per message in my session) the auto-compact threshold matters more: five heavy messages cost $56 at 950K, $38 at 600K, $32 at 400K, and $26 at 200K.

Related: #8144 (the banner), #11999 (moves it out of the composer stack), #7590 (cache countdown idea).

Impact

Major degradation or frequent failure

Version or commit

0.0.43-nightly.20260926.2282 (code checked on main @ de251fc)

Environment

Desktop app on Linux (NixOS, Wayland, niri); Claude Code 2.1.283 with Opus 5.5; Codex 0.157.1 with gpt-6-astra

Logs or stack traces

Screenshots, recordings, or supporting files

No response

Workaround

Compact by hand within an hour of the last reply, or lower "Auto-compact after" in the Claude provider settings.

Activity

  1. added
    bugSomething is broken or behaving incorrectly.
    needs-triageIssue needs maintainer review and initial categorization.
    on Sep 27, 2026
  2. juliusmarminge commented on Sep 27, 2026

    @juliusmarminge
    Member

    Triage

    The 70-minute banner is the behavior shipped in #8144, and the cost gap you measured is real.

    shouldOfferResumeCompaction (apps/web/src/components/chat/ContextWindowMeter.logic.ts) shows Resume with less context only for claudeAgent, only at 100,000 tokens or more, and only when the latest context-window.updated activity is at least 70 minutes old. That timestamp is the clock. Tests lock the 70-minute / 100k pair to Claude Code's resume rule. Compact then sends /compact as a new turn (ChatView.tsx).

    Turn usage keeps a single cache_creation_input_tokens total (normalizeClaudeTurnTokenUsage in ClaudeAdapter.ts). The ephemeral_1h_input_tokens / ephemeral_5m_input_tokens split is discarded, so the banner has no cache TTL to race. It always appears after both a 60-minute and a 5-minute cache have expired, and the compact pays the cold read. Desktop notifications only fire for turn completion, input, approval, and failure while a client is connected. Nothing notifies or compacts during the warm window.

    Codex is left out of that function on purpose. It can already compact: the provider advertises /compact, and CodexAdapter.compactThread calls thread/compact/start. There is no idle offer.

    Claude's own resume_return dialog is a second path. The adapter only relays it (handleResumeDialog). Its timing and the tengu_* experiments stay on Claude Code's side.

    Auto-compact after is already any integer from 100,000 to 1,000,000 (ClaudeSettings.autoCompactWindow). 400000 and 600000 can be typed today. There is no preset list and no separate "suggest compaction from N tokens" setting.

    #7590 does not resolve to an issue or pull request. #11999 moves this notice in the UI and leaves the timer alone.

    The part that matches the existing feature is a warm-cache offer: compact, or notify, before the cache window closes, using the 1-hour vs 5-minute split when the CLI reports it. A lower banner threshold only helps while the app is open inside that window. "Continue in a new thread" (brief, link, archive after the first reply) and a Codex idle offer are separate features.

    Checked on main @ de251fc.

  3. added
    enhancementRequested improvement or new capability.
    via-triageFiled through npx t3 triage
    and removed
    bugSomething is broken or behaving incorrectly.
    needs-triageIssue needs maintainer review and initial categorization.
    on Sep 27, 2026
  4. Vantrongs commented on Sep 29, 2026

    @Vantrongs
    ContributorAuthor

    Looks like the triage script has a small blind spot with discussions. In the triage above, #7590 came out as "does not resolve to an issue or pull request", but it's a discussion: #7590

    Likely cause: gh issue view and GraphQL issueOrPullRequest only know issues and PRs, so a discussion number comes back as not found. Checking repository.discussion(number:) when that happens should fix it.

  5. MartinPTielemans commented on Oct 7, 2026

    @MartinPTielemans

    @juliusmarminge Following up on your triage above ("compact, or notify, before the cache window closes"): I've put up a narrow slice of that as draft #16855 and would like a yes/no on direction before going further.

    It adds an opt-in compactBeforeWaiting flag to watch_pull_request. When the user asks for it, starting a watch queues one native /compact right after the current turn, while the cache is most likely still warm, instead of paying a cold read when PR news arrives hours later. It doesn't time the cache TTL, send keepalives, or change the existing resume banner; ordinary watches are unchanged. It works for any provider that implements native compaction.

    If this direction works for you, I'll link your approval and mark it ready. If you'd rather this go through the TTL-aware offer instead, I'm happy to close it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementRequested improvement or new capability.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions