Repository navigation
🔎 feat: Per-Search Model-Output Budget for Web Search Highlights - #242
Conversation
Cap the highlight content the web search tool feeds the model per search via a new `maxOutputChars` config (default 50,000 chars; also settable through `SEARCH_MAX_LLM_OUTPUT_CHARS`). Previously `formatResultsForLLM` aggregated every source's reranked highlights with no aggregate budget, so a 5-source search over 50K-char scrapes could dump ~150-190K tokens into a single turn and trigger premature summarization/compaction. Highlights are now walked in relevance order (organic, then news) and kept whole until the budget is hit; the boundary highlight is truncated when meaningful room remains, otherwise dropped, with an omission marker appended. Snippets/titles/URLs are untouched and full content always remains in the `WEB_SEARCH` artifact for citations.
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5083179bc0
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| } | ||
| const remaining = maxChars - used; | ||
| if (remaining >= MIN_PARTIAL_HIGHLIGHT_CHARS) { | ||
| kept.push({ ...highlight, text: `${highlight.text.slice(0, remaining)}\n…[truncated]` }); |
There was a problem hiding this comment.
Recompute references after truncating highlights
When the budget boundary falls before or inside a citation marker in a highlighted excerpt, this keeps the original highlight.references while replacing text with a raw slice. formatSource later emits Core References from that stale array, so the model can see references for citations that are no longer present in the visible excerpt, or a broken partial marker. Truncate on citation-safe boundaries or rebuild the references from the sliced text before formatting.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 875a081. The truncated boundary highlight is now rebuilt as { score, text } only, dropping references, so formatSource no longer emits Core References (or anchors) for citations that fall in the cut tail. Whole (untruncated) highlights keep their references intact. Added tests for both cases.
| } | ||
| trimmed++; | ||
| } | ||
| source.highlights = kept; |
There was a problem hiding this comment.
Preserve untrimmed highlights for artifacts
When a search exceeds maxOutputChars, this writes the budgeted highlight list back onto the same searchResult object that createTool later spreads into the WEB_SEARCH artifact after formatting. As a result, the excerpt text omitted from the model-facing output is not preserved for UI/citation recovery despite the new contract that only the prompt text is bounded; build the model-facing copy separately and leave the original result data intact.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Looked into this and I don't think it's a regression. Highlights are never carried in the WEB_SEARCH artifact: formatResultsForLLM already deletes results.organic[i].highlights (and the topStories equivalent) after formatting, in both the old and new code paths. The per-source content field — the actual full scraped text used for recovery — is untouched by the budget.
For references: they're intentionally collected only from the highlights that are actually rendered, because the in-text turn{t}ref{N} anchors index into the same references array by position. Preserving omitted highlights' references would add entries with no corresponding anchor and shift the indices, breaking the anchor→reference mapping for the citations the model did see. So trimming references alongside highlights is what keeps the artifact consistent with the model-facing output.
Happy to revisit if the goal is to surface full highlights in the artifact UI, but that would be changing the pre-existing post-format delete behavior, which is outside this PR's scope.
| } | ||
| const kept: t.Highlight[] = []; | ||
| for (const highlight of highlights) { | ||
| const length = highlight.text.length; |
There was a problem hiding this comment.
Ignore blank highlights when spending budget
If a reranker returns a blank or whitespace-only highlight before useful ones, this charges its raw length against maxOutputChars, but formatSource later filters blank highlights with h.text.trim().length > 0. In that scenario the budget can be exhausted by text that is never shown, causing real highlights to be omitted or leaving only a truncated blank marker; skip or normalize empty highlights before counting them.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 875a081. trimHighlightsToBudget now trims each highlight's text up front and skips empty/whitespace-only ones before charging the budget — matching formatSource's text.trim().length > 0 filter — so blanks can no longer exhaust the budget ahead of real highlights. The budget is also charged on the trimmed length to match what's actually rendered. Added a test.
…ight refs Address two Codex findings on the per-search budget: - Skip blank/whitespace-only highlights before charging the budget. They are filtered out by `formatSource` (`text.trim().length > 0`) and never rendered, so charging their raw length could starve real highlights. - Drop a truncated boundary highlight's references. Slicing the text can cut past a citation marker, so keeping the original `references` would emit Core References for citations no longer visible to the model. Budget is now charged on trimmed length to match what is rendered.
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 875a081c0b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| } | ||
| const remaining = maxChars - used; | ||
| if (remaining >= MIN_PARTIAL_HIGHLIGHT_CHARS) { | ||
| kept.push({ score: highlight.score, text: `${text.slice(0, remaining)}\n…[truncated]` }); |
There was a problem hiding this comment.
Preserve visible citations in truncated highlights
When the boundary highlight is truncated and the kept prefix still contains a citation marker, this replacement drops highlight.references, so formatSource cannot emit the corresponding Core References and the returned references array omits citations that are still visible to the model. This affects searches where a cited passage crosses the per-search budget; consider filtering/retracking references against the sliced text instead of discarding all of them.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in e0b0cf9. Truncation now filters highlight.references to those whose (type#N) marker still appears in the kept prefix — matched with the same marker regex highlights.ts uses to build them — rather than discarding all of them. Citations visible in the prefix are retained (with their Core References), and only those whose marker was cut from the tail are dropped. Added a test with one marker in the prefix and one in the tail.
| if (remaining >= MIN_PARTIAL_HIGHLIGHT_CHARS) { | ||
| kept.push({ score: highlight.score, text: `${text.slice(0, remaining)}\n…[truncated]` }); | ||
| used = maxChars; | ||
| } | ||
| trimmed++; |
There was a problem hiding this comment.
Stop after dropping the boundary highlight
When a high-relevance highlight exceeds the remaining budget with less than MIN_PARTIAL_HIGHLIGHT_CHARS left, this branch drops it but leaves used below maxChars and continues iterating, so a later lower-ranked short highlight can still be emitted. This breaks the relevance-ordered budget policy for searches with one long boundary highlight followed by short ones; mark the budget exhausted or stop scanning once the boundary highlight is dropped.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in e0b0cf9. used is now set to maxChars whenever the boundary highlight can't be kept whole — in the drop branch too, not just on truncation — so once relevance order forces a drop, every later (lower-ranked) highlight is dropped as well. Added a test: a long boundary highlight with <MIN room followed by a short one now correctly omits both.
…stop - Truncated highlights now keep references whose `(type#N)` marker survives in the kept prefix (matched via the same marker regex as highlights.ts), instead of dropping all references. Preserves still-visible citations while still discarding those cut from the tail. - Mark the budget exhausted when the boundary highlight is dropped for having less than MIN_PARTIAL room, not just when it is truncated. Stops a later lower-ranked short highlight from slipping in past a dropped higher-ranked one, keeping the relevance-ordered prefix intact.
|
@codex review |
|
Codex Review: Didn't find any major issues. Already looking forward to the next diff. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
Summary
Adds a per-search budget on the model-facing highlight content the web search tool emits, so a single search can no longer flood the context window and trigger premature summarization/compaction.
formatResultsForLLMpreviously aggregated every source's reranked highlights with no aggregate cap. A typical 5-source search over 50K-char scrapes could push ~150-190K tokens into one turn — enough to blow past the context budget on its own.What changed
maxOutputCharsconfig onSearchToolConfig(threadedcreateSearchTool → createTool → formatResultsForLLM).SEARCH_MAX_LLM_OUTPUT_CHARSenv var →50000char default.maxContentLength(which caps scraped/reranked content per source). Full content always remains in theWEB_SEARCHartifact for citations/UI — only the model-facing prompt text is bounded.trimHighlightsToBudgetwalks sources in relevance order (organic first, then news; highlights in their reranked order), keeping whole highlights until the budget is hit. The boundary highlight is truncated when meaningful room remains (≥200 chars), otherwise dropped.Why absolute chars (not a %)
The tool is built at agent-init, before any live "remaining context" is known, and budgeting against remaining context would starve search instead of summarizing. So the SDK takes a dumb absolute char budget; a host that knows its context window (e.g. LibreChat) computes a window-relative value and passes it in. Default keeps standalone consumers bounded.
Tests
src/tools/search/format.test.ts(10 cases): budget enforcement, relevance-ordered trim, boundary truncation vs. drop, snippets/titles always kept, omission marker presence/absence, andresolveMaxLLMOutputCharsconfig/env/default precedence.