Skip to content

Support disabling thinking mode + client-side context compaction for DeepSeek (OpenAI-compatible models without server-side compaction) #11037

Description

@OntologyAgent

Description

Summary

Using deepseek-flash (DeepSeek-V4.1, 1M context) as the marimo AI assistant model causes "context length exceeded" errors after only a few multi-turn conversations. This is not a raw input-size problem — it is the combination of DeepSeek's always-on thinking mode and the absence of any context compaction for this provider.

Problem

  • Configure marimo AI assistant with deepseek/deepseek-flash via base_url = https://api.deepseek.com.
  • A conversation that would be tiny in raw tokens (a 12k-token notebook context) blows past the 1M context window after several assistant turns involving tool calls.
  • The same workload does not fail this way with OpenAI/Anthropic models.

Root cause

There are three compounding issues:

1. DeepSeek thinking mode defaults ON (high effort) and marimo cannot turn it off

DeepSeek's official docs state that thinking mode is enabled by default with effort = high:

marimo's OpenAIProvider._default_thinking deliberately returns None for custom base_url values (so no thinking/reasoning_effort is sent):

  • marimo/_server/ai/providers.py

For OpenAI this is fine, but for DeepSeek "not sending a flag" means "fall back to the server default", which is thinking ON at high effort.

There is also no user-facing way to override this: neither AiConfig nor OpenAiConfig in marimo/_config/config.py exposes thinking, reasoning_effort, or an extra_body passthrough. So a user cannot send DeepSeek's documented switch extra_body={"thinking": {"type": "disabled"}}.

2. DeepSeek accumulates reasoning_content into context when tools are present

DeepSeek's docs state that when a request carries tools, the model's reasoning_content from prior turns must be sent back and is concatenated into the context:

marimo's assistant is a multi-turn agent that always uses tools, so the high-effort reasoning content from every turn is retained and grows without bound.

3. No context compaction exists for DeepSeek

pydantic-ai's compaction (CompactionPart) is provider-initiated — it only works when the provider (OpenAI Responses API / Anthropic) returns a server-side compaction item. The Agent constructor has no client-side compaction parameters.

DeepSeek has no server-side compaction. Its only related feature is the KV-cache disk cache, which is a Sliding Window Attention cache for cost/latency, not a context-size reduction:

marimo passes the full message history straight through (agent.run_stream(message_history=...) in marimo/_ai/llm/_impl.py), so nothing ever trims the accumulated history.

Proposed solution

Any of the following (ideally 1 + 2 together):

  1. Expose thinking-mode control for OpenAI-compatible providers. Add a thinking / reasoning_effort / extra_body field to OpenAiConfig (and wire it into the pydantic-ai model settings), so users can send {"thinking": {"type": "disabled"}} to DeepSeek.

  2. Add client-side context compaction for providers without server-side compaction. When the message history approaches the model's token limit, summarize older messages (or drop/truncate them) before sending. This would also help other OpenAI-compatible providers (Ollama, vLLM, local models).

  3. (Minor) Refresh the DeepSeek model recognition. pydantic-ai 2.47.0's DeepSeekModelName only knows deepseek-v4-flash / deepseek-v4-pro, not the current deepseek-flash model name. The profile's is_v4 = model_name.startswith('deepseek-v4-') check misses deepseek-flash, so thinking support is misdetected. Bumping/updating the DeepSeek profile would help.

Current workaround

Keep assistant sessions short and start a new session frequently, since context growth is per-session accumulation.

Environment

  • marimo: 0.24.2
  • pydantic-ai: 2.47.0
  • Model: deepseek-flash (DeepSeek-V4.1, 1M context / 384K max output)

Are you willing to submit a PR? (You must receive approval from the team before submitting a PR.)

  • Yes

Alternatives

No response

Additional context

No response

Activity

  1. Light2Dark commented on Sep 30, 2026

    @Light2Dark
    Member

    Thanks for the report! This is partially an upstream issue and partially ours as well.
    pydantic-ai at the moment doesn't translate deepseek-flash's profile correctly pydantic/pydantic-ai#8353.

    Some questions though:

    1. How are you using the AI, through chat sidebar? Or other ways like Generate with AI. Because if you'd like to control thinking, are you imagining controlling it at the config level, or at a conversation level?
    2. I have some plans to add better compaction, history processing for the chat sidebar. Still early works experiment(ai): benchmark a more reliable code-mode harness #11030
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions