Description
Summary
Using deepseek-flash (DeepSeek-V4.1, 1M context) as the marimo AI assistant model causes "context length exceeded" errors after only a few multi-turn conversations. This is not a raw input-size problem — it is the combination of DeepSeek's always-on thinking mode and the absence of any context compaction for this provider.
Problem
- Configure marimo AI assistant with
deepseek/deepseek-flash via base_url = https://api.deepseek.com.
- A conversation that would be tiny in raw tokens (a 12k-token notebook context) blows past the 1M context window after several assistant turns involving tool calls.
- The same workload does not fail this way with OpenAI/Anthropic models.
Root cause
There are three compounding issues:
1. DeepSeek thinking mode defaults ON (high effort) and marimo cannot turn it off
DeepSeek's official docs state that thinking mode is enabled by default with effort = high:
marimo's OpenAIProvider._default_thinking deliberately returns None for custom base_url values (so no thinking/reasoning_effort is sent):
marimo/_server/ai/providers.py
For OpenAI this is fine, but for DeepSeek "not sending a flag" means "fall back to the server default", which is thinking ON at high effort.
There is also no user-facing way to override this: neither AiConfig nor OpenAiConfig in marimo/_config/config.py exposes thinking, reasoning_effort, or an extra_body passthrough. So a user cannot send DeepSeek's documented switch extra_body={"thinking": {"type": "disabled"}}.
2. DeepSeek accumulates reasoning_content into context when tools are present
DeepSeek's docs state that when a request carries tools, the model's reasoning_content from prior turns must be sent back and is concatenated into the context:
marimo's assistant is a multi-turn agent that always uses tools, so the high-effort reasoning content from every turn is retained and grows without bound.
3. No context compaction exists for DeepSeek
pydantic-ai's compaction (CompactionPart) is provider-initiated — it only works when the provider (OpenAI Responses API / Anthropic) returns a server-side compaction item. The Agent constructor has no client-side compaction parameters.
DeepSeek has no server-side compaction. Its only related feature is the KV-cache disk cache, which is a Sliding Window Attention cache for cost/latency, not a context-size reduction:
marimo passes the full message history straight through (agent.run_stream(message_history=...) in marimo/_ai/llm/_impl.py), so nothing ever trims the accumulated history.
Proposed solution
Any of the following (ideally 1 + 2 together):
-
Expose thinking-mode control for OpenAI-compatible providers. Add a thinking / reasoning_effort / extra_body field to OpenAiConfig (and wire it into the pydantic-ai model settings), so users can send {"thinking": {"type": "disabled"}} to DeepSeek.
-
Add client-side context compaction for providers without server-side compaction. When the message history approaches the model's token limit, summarize older messages (or drop/truncate them) before sending. This would also help other OpenAI-compatible providers (Ollama, vLLM, local models).
-
(Minor) Refresh the DeepSeek model recognition. pydantic-ai 2.47.0's DeepSeekModelName only knows deepseek-v4-flash / deepseek-v4-pro, not the current deepseek-flash model name. The profile's is_v4 = model_name.startswith('deepseek-v4-') check misses deepseek-flash, so thinking support is misdetected. Bumping/updating the DeepSeek profile would help.
Current workaround
Keep assistant sessions short and start a new session frequently, since context growth is per-session accumulation.
Environment
- marimo: 0.24.2
- pydantic-ai: 2.47.0
- Model: deepseek-flash (DeepSeek-V4.1, 1M context / 384K max output)
Are you willing to submit a PR? (You must receive approval from the team before submitting a PR.)
Alternatives
No response
Additional context
No response
Description
Summary
Using
deepseek-flash(DeepSeek-V4.1, 1M context) as the marimo AI assistant model causes "context length exceeded" errors after only a few multi-turn conversations. This is not a raw input-size problem — it is the combination of DeepSeek's always-on thinking mode and the absence of any context compaction for this provider.Problem
deepseek/deepseek-flashviabase_url = https://api.deepseek.com.Root cause
There are three compounding issues:
1. DeepSeek thinking mode defaults ON (high effort) and marimo cannot turn it off
DeepSeek's official docs state that thinking mode is enabled by default with
effort = high:marimo's
OpenAIProvider._default_thinkingdeliberately returnsNonefor custombase_urlvalues (so nothinking/reasoning_effortis sent):marimo/_server/ai/providers.pyFor OpenAI this is fine, but for DeepSeek "not sending a flag" means "fall back to the server default", which is thinking ON at high effort.
There is also no user-facing way to override this: neither
AiConfignorOpenAiConfiginmarimo/_config/config.pyexposesthinking,reasoning_effort, or anextra_bodypassthrough. So a user cannot send DeepSeek's documented switchextra_body={"thinking": {"type": "disabled"}}.2. DeepSeek accumulates
reasoning_contentinto context when tools are presentDeepSeek's docs state that when a request carries
tools, the model'sreasoning_contentfrom prior turns must be sent back and is concatenated into the context:marimo's assistant is a multi-turn agent that always uses tools, so the high-effort reasoning content from every turn is retained and grows without bound.
3. No context compaction exists for DeepSeek
pydantic-ai's compaction (
CompactionPart) is provider-initiated — it only works when the provider (OpenAI Responses API / Anthropic) returns a server-side compaction item. TheAgentconstructor has no client-side compaction parameters.DeepSeek has no server-side compaction. Its only related feature is the KV-cache disk cache, which is a Sliding Window Attention cache for cost/latency, not a context-size reduction:
marimo passes the full message history straight through (
agent.run_stream(message_history=...)inmarimo/_ai/llm/_impl.py), so nothing ever trims the accumulated history.Proposed solution
Any of the following (ideally 1 + 2 together):
Expose thinking-mode control for OpenAI-compatible providers. Add a
thinking/reasoning_effort/extra_bodyfield toOpenAiConfig(and wire it into the pydantic-ai model settings), so users can send{"thinking": {"type": "disabled"}}to DeepSeek.Add client-side context compaction for providers without server-side compaction. When the message history approaches the model's token limit, summarize older messages (or drop/truncate them) before sending. This would also help other OpenAI-compatible providers (Ollama, vLLM, local models).
(Minor) Refresh the DeepSeek model recognition. pydantic-ai 2.47.0's
DeepSeekModelNameonly knowsdeepseek-v4-flash/deepseek-v4-pro, not the currentdeepseek-flashmodel name. The profile'sis_v4 = model_name.startswith('deepseek-v4-')check missesdeepseek-flash, so thinking support is misdetected. Bumping/updating the DeepSeek profile would help.Current workaround
Keep assistant sessions short and start a new session frequently, since context growth is per-session accumulation.
Environment
Are you willing to submit a PR? (You must receive approval from the team before submitting a PR.)
Alternatives
No response
Additional context
No response