Skip to content

🔎 feat: Per-Search Model-Output Budget for Web Search Highlights - #242

Merged
danny-avila merged 3 commits into
mainfrom
claude/search-output-budget
Jun 16, 2026
Merged

danny-avila merged 3 commits into
mainfrom
claude/search-output-budget

Conversation

@danny-avila

Copy link
Copy Markdown
Collaborator

Summary

Adds a per-search budget on the model-facing highlight content the web search tool emits, so a single search can no longer flood the context window and trigger premature summarization/compaction.

formatResultsForLLM previously aggregated every source's reranked highlights with no aggregate cap. A typical 5-source search over 50K-char scrapes could push ~150-190K tokens into one turn — enough to blow past the context budget on its own.

What changed

  • New maxOutputChars config on SearchToolConfig (threaded createSearchTool → createTool → formatResultsForLLM).
    • Resolution order: explicit config value → SEARCH_MAX_LLM_OUTPUT_CHARS env var → 50000 char default.
    • Distinct from the existing maxContentLength (which caps scraped/reranked content per source). Full content always remains in the WEB_SEARCH artifact for citations/UI — only the model-facing prompt text is bounded.
  • trimHighlightsToBudget walks sources in relevance order (organic first, then news; highlights in their reranked order), keeping whole highlights until the budget is hit. The boundary highlight is truncated when meaningful room remains (≥200 chars), otherwise dropped.
  • Omission marker appended when anything was trimmed, pointing the model at the cited sources for full content.
  • Snippets / titles / URLs are never touched (small, high-signal).

Why absolute chars (not a %)

The tool is built at agent-init, before any live "remaining context" is known, and budgeting against remaining context would starve search instead of summarizing. So the SDK takes a dumb absolute char budget; a host that knows its context window (e.g. LibreChat) computes a window-relative value and passes it in. Default keeps standalone consumers bounded.

Tests

  • New src/tools/search/format.test.ts (10 cases): budget enforcement, relevance-ordered trim, boundary truncation vs. drop, snippets/titles always kept, omission marker presence/absence, and resolveMaxLLMOutputChars config/env/default precedence.
  • Full search suite green (71 tests), typecheck clean, lint clean.

Cap the highlight content the web search tool feeds the model per search
via a new `maxOutputChars` config (default 50,000 chars; also settable
through `SEARCH_MAX_LLM_OUTPUT_CHARS`). Previously `formatResultsForLLM`
aggregated every source's reranked highlights with no aggregate budget,
so a 5-source search over 50K-char scrapes could dump ~150-190K tokens
into a single turn and trigger premature summarization/compaction.

Highlights are now walked in relevance order (organic, then news) and
kept whole until the budget is hit; the boundary highlight is truncated
when meaningful room remains, otherwise dropped, with an omission marker
appended. Snippets/titles/URLs are untouched and full content always
remains in the `WEB_SEARCH` artifact for citations.
@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5083179bc0

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/tools/search/format.ts Outdated
}
const remaining = maxChars - used;
if (remaining >= MIN_PARTIAL_HIGHLIGHT_CHARS) {
kept.push({ ...highlight, text: `${highlight.text.slice(0, remaining)}\n…[truncated]` });

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Recompute references after truncating highlights

When the budget boundary falls before or inside a citation marker in a highlighted excerpt, this keeps the original highlight.references while replacing text with a raw slice. formatSource later emits Core References from that stale array, so the model can see references for citations that are no longer present in the visible excerpt, or a broken partial marker. Truncate on citation-safe boundaries or rebuild the references from the sliced text before formatting.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 875a081. The truncated boundary highlight is now rebuilt as { score, text } only, dropping references, so formatSource no longer emits Core References (or anchors) for citations that fall in the cut tail. Whole (untruncated) highlights keep their references intact. Added tests for both cases.

}
trimmed++;
}
source.highlights = kept;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve untrimmed highlights for artifacts

When a search exceeds maxOutputChars, this writes the budgeted highlight list back onto the same searchResult object that createTool later spreads into the WEB_SEARCH artifact after formatting. As a result, the excerpt text omitted from the model-facing output is not preserved for UI/citation recovery despite the new contract that only the prompt text is bounded; build the model-facing copy separately and leave the original result data intact.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looked into this and I don't think it's a regression. Highlights are never carried in the WEB_SEARCH artifact: formatResultsForLLM already deletes results.organic[i].highlights (and the topStories equivalent) after formatting, in both the old and new code paths. The per-source content field — the actual full scraped text used for recovery — is untouched by the budget.

For references: they're intentionally collected only from the highlights that are actually rendered, because the in-text turn{t}ref{N} anchors index into the same references array by position. Preserving omitted highlights' references would add entries with no corresponding anchor and shift the indices, breaking the anchor→reference mapping for the citations the model did see. So trimming references alongside highlights is what keeps the artifact consistent with the model-facing output.

Happy to revisit if the goal is to surface full highlights in the artifact UI, but that would be changing the pre-existing post-format delete behavior, which is outside this PR's scope.

Comment thread src/tools/search/format.ts Outdated
}
const kept: t.Highlight[] = [];
for (const highlight of highlights) {
const length = highlight.text.length;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Ignore blank highlights when spending budget

If a reranker returns a blank or whitespace-only highlight before useful ones, this charges its raw length against maxOutputChars, but formatSource later filters blank highlights with h.text.trim().length > 0. In that scenario the budget can be exhausted by text that is never shown, causing real highlights to be omitted or leaving only a truncated blank marker; skip or normalize empty highlights before counting them.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 875a081. trimHighlightsToBudget now trims each highlight's text up front and skips empty/whitespace-only ones before charging the budget — matching formatSource's text.trim().length > 0 filter — so blanks can no longer exhaust the budget ahead of real highlights. The budget is also charged on the trimmed length to match what's actually rendered. Added a test.

…ight refs

Address two Codex findings on the per-search budget:

- Skip blank/whitespace-only highlights before charging the budget. They
  are filtered out by `formatSource` (`text.trim().length > 0`) and never
  rendered, so charging their raw length could starve real highlights.
- Drop a truncated boundary highlight's references. Slicing the text can
  cut past a citation marker, so keeping the original `references` would
  emit Core References for citations no longer visible to the model.

Budget is now charged on trimmed length to match what is rendered.
@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 875a081c0b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/tools/search/format.ts Outdated
}
const remaining = maxChars - used;
if (remaining >= MIN_PARTIAL_HIGHLIGHT_CHARS) {
kept.push({ score: highlight.score, text: `${text.slice(0, remaining)}\n…[truncated]` });

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve visible citations in truncated highlights

When the boundary highlight is truncated and the kept prefix still contains a citation marker, this replacement drops highlight.references, so formatSource cannot emit the corresponding Core References and the returned references array omits citations that are still visible to the model. This affects searches where a cited passage crosses the per-search budget; consider filtering/retracking references against the sliced text instead of discarding all of them.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in e0b0cf9. Truncation now filters highlight.references to those whose (type#N) marker still appears in the kept prefix — matched with the same marker regex highlights.ts uses to build them — rather than discarding all of them. Citations visible in the prefix are retained (with their Core References), and only those whose marker was cut from the tail are dropped. Added a test with one marker in the prefix and one in the tail.

Comment on lines +62 to +66
if (remaining >= MIN_PARTIAL_HIGHLIGHT_CHARS) {
kept.push({ score: highlight.score, text: `${text.slice(0, remaining)}\n…[truncated]` });
used = maxChars;
}
trimmed++;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Stop after dropping the boundary highlight

When a high-relevance highlight exceeds the remaining budget with less than MIN_PARTIAL_HIGHLIGHT_CHARS left, this branch drops it but leaves used below maxChars and continues iterating, so a later lower-ranked short highlight can still be emitted. This breaks the relevance-ordered budget policy for searches with one long boundary highlight followed by short ones; mark the budget exhausted or stop scanning once the boundary highlight is dropped.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in e0b0cf9. used is now set to maxChars whenever the boundary highlight can't be kept whole — in the drop branch too, not just on truncation — so once relevance order forces a drop, every later (lower-ranked) highlight is dropped as well. Added a test: a long boundary highlight with <MIN room followed by a short one now correctly omits both.

…stop

- Truncated highlights now keep references whose `(type#N)` marker survives
  in the kept prefix (matched via the same marker regex as highlights.ts),
  instead of dropping all references. Preserves still-visible citations
  while still discarding those cut from the tail.
- Mark the budget exhausted when the boundary highlight is dropped for
  having less than MIN_PARTIAL room, not just when it is truncated. Stops a
  later lower-ranked short highlight from slipping in past a dropped
  higher-ranked one, keeping the relevance-ordered prefix intact.
@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Already looking forward to the next diff.

Reviewed commit: e0b0cf95a6

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@danny-avila
danny-avila merged commit 567fa06 into main Jun 16, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant