Repository navigation
Compaction session compacts a pending approval call, leaving an orphan output #5340
Description
Activity
Independent extension of the reproduction, with AI assistance: the rejection path also leaves an unpaired output on the next turn, even though the protected tool never executes. I checked both
Runner.runand fully consumedRunner.run_streamedafter aRunState.to_json()/RunState.from_json()round trip.Environment: released
openai-agents==0.23.1, Python 3.12.11,openai==3.28.0, Pydantic 2.14.0. Public runner/session APIs, real SQLite,ScriptedModel, and a syntheticresponses.compactresponse; no provider calls.Decision Automatic compaction hook Tool dispatches Unpaired pending-c1output in next-turn model inputApprove Always true 1 Present in both runner modes Reject Always true 0 Present in both runner modes Approve Always false 1 Absent in both runner modes Reject Always false 0 Absent in both runner modes Eight cases executed: four next-turn pairing violations and four controls. In all eight, the immediate resumed model input still has its call/output pairing; the missing call is observable in the persisted history and the following user turn. Rejection remains enforced: this is not a rejected-tool execution.
For the regression, I would include rejection as well as approval: a denied call still produces a tool output that needs its original call. Deferring only when a backend actually executes would miss that case. The JSON round trip keeps this tied to the public durable state boundary.
Scope: one process, a stub compaction result and a scripted model. No real compaction API, OpenAI 400 response, separate-worker restore, SDK patch, or full-suite claim. The false compaction hook is a comparison control, not a complete fix.
Runnable eight-case characterization
Save as
characterize_approval_compaction.pyand run:uv run --no-project --python 3.12 --with openai-agents==0.23.1 python characterize_approval_compaction.py
from __future__ import annotations import asyncio import importlib.metadata import json import platform import tempfile from pathlib import Path from types import SimpleNamespace from typing import Any from agents import Agent, Runner, RunState, SQLiteSession, function_tool from agents.memory import OpenAIResponsesCompactionSession from agents.testing import ScriptedModel, assistant_message, function_call from agents.tracing import set_tracing_disabled set_tracing_disabled(True) def normalized(items: list[Any]) -> list[dict[str, Any]]: return [ item if isinstance(item, dict) else item.model_dump(exclude_unset=True) for item in items ] def unpaired_outputs(items: list[Any]) -> list[str]: values = normalized(items) calls = {item["call_id"] for item in values if item.get("type") == "function_call"} return [ item["call_id"] for item in values if item.get("type") == "function_call_output" and item["call_id"] not in calls ] async def run_case(*, streaming: bool, allow: bool, compact: bool) -> dict[str, Any]: dispatches: list[str] = [] compact_inputs: list[list[dict[str, Any]]] = [] @function_tool(needs_approval=True) def protected(value: str) -> str: """Return a synthetic value after approval.""" dispatches.append(value) return "synthetic:" + value async def fake_compact(**kwargs: Any) -> SimpleNamespace: compact_inputs.append(normalized(kwargs.get("input", []))) return SimpleNamespace( output=[ { "type": "compaction", "id": "synthetic-compaction", "encrypted_content": "synthetic-summary", } ], usage=None, ) model = ScriptedModel( steps=[ [function_call("protected", {"value": "v"}, call_id="pending-c1")], [assistant_message("finished")], [assistant_message("next answer")], ] ) agent = Agent(name="synthetic-agent", model=model, tools=[protected]) async def run(value: str | RunState, session: Any) -> Any: if not streaming: return await Runner.run(agent, value, session=session) result = Runner.run_streamed(agent, value, session=session) async for _ in result.stream_events(): pass return result with tempfile.TemporaryDirectory(prefix="synthetic-compaction-") as directory: store = SQLiteSession("synthetic-session", str(Path(directory) / "history.db")) try: session = OpenAIResponsesCompactionSession( "synthetic-session", store, client=SimpleNamespace(responses=SimpleNamespace(compact=fake_compact)), compaction_mode="input", should_trigger_compaction=lambda _: compact, ) first = await run("first question", session) assert len(first.interruptions) == 1 and not dispatches stored_at_pause = normalized(await session.get_items()) state = first.to_state() state_json = state.to_json() restored = await RunState.from_json(agent, state_json) interruption = first.interruptions[0] if allow: restored.approve(interruption) else: restored.reject(interruption) await run(restored, session) stored_after_resume = normalized(await session.get_items()) resume_model_input = normalized(model.calls[-1].input) await run("next question", session) next_model_input = normalized(model.calls[-1].input) assert len(dispatches) == (1 if allow else 0) return { "streaming": streaming, "decision": "approve" if allow else "reject", "compaction_enabled": compact, "runstate_json_roundtrip": True, "dispatches": len(dispatches), "compact_invocations": len(compact_inputs), "stored_at_pause": stored_at_pause, "stored_after_resume": stored_after_resume, "resume_model_input": resume_model_input, "next_model_input": next_model_input, "resume_unpaired_outputs": unpaired_outputs(resume_model_input), "next_turn_unpaired_outputs": unpaired_outputs(next_model_input), } finally: store.close() async def main() -> None: cases = [] for streaming in (False, True): for allow in (True, False): for compact in (True, False): cases.append( await run_case(streaming=streaming, allow=allow, compact=compact) ) violations = [case for case in cases if case["next_turn_unpaired_outputs"]] report = { "python": platform.python_version(), "openai_agents": importlib.metadata.version("openai-agents"), "openai": importlib.metadata.version("openai"), "pydantic": importlib.metadata.version("pydantic"), "cases_executed": len(cases), "cases_with_next_turn_unpaired_outputs": len(violations), "cases": cases, "scope": ( "Public SDK Runner, real SQLite session and RunState JSON round trip; " "ScriptedModel and synthetic compaction response, one process. " "No model/compaction API calls, separate-process restore, SDK fix or full suite." ), } Path(__file__).with_name("verification.json").write_text( json.dumps(report, indent=2) + "\n" ) print( json.dumps( { **{key: value for key, value in report.items() if key != "cases"}, "case_summaries": [ { key: case[key] for key in ( "streaming", "decision", "compaction_enabled", "dispatches", "resume_unpaired_outputs", "next_turn_unpaired_outputs", ) } for case in cases ], }, indent=2, ) ) if __name__ == "__main__": asyncio.run(main())
Please read this first
function_call_outputon a different path (deferred session items on resume). Compaction unusable in Runner.run_streamed: responses.compact runs before tool outputs are in chain #2317 was about compaction running before tool outputs in streamed runs. I found nothing for compaction at an approval interruption.Describe the bug
With
OpenAIResponsesCompactionSession, automatic compaction runs right after a turn that ends in a tool approval interruption. That turn stored afunction_callwith no output yet, so the pending call is sent toresponses.compactand replaced by the compaction item. After the approval, the resumed run appends thefunction_call_output, whosefunction_callis no longer in the session, and every later run sends that unpaired output to the model.prepare_input_with_sessiononly prunes orphan outputs whenSessionSettings.limitis set, so on a session without a limit nothing removes it.Debug information
mainat 125efa0, unchanged)any-llm,litellm, orpydantic): none neededagents.testing.ScriptedModeland a stubresponses.compactclient, no networkOutput of the script below on 0.23.1:
Repro steps
The compaction client is a stub that returns one
compactionitem; nothing calls the API.Expected behavior
A
function_callthat is waiting for approval should not be compacted away from the output it is about to receive. Either compaction is deferred until the resumed turn has stored the output (as already happens for turns with local tool outputs), or the call and its output stay together, so nofunction_call_output(c1)is stored or sent withoutfunction_call(c1).Cause.
_apply_post_write_compaction(src/agents/run_internal/session_persistence.py701-714 at 125efa0) defers automatic compaction only when the saved batch has local tool or handoff outputs. A turn that ends in an interruption has pending calls and no outputs, so it falls through to an immediate compaction. InOpenAIResponsesCompactionSession(src/agents/memory/openai_responses_compaction_session.py397-415) the partial-suffix branch refuses to compact whendrop_orphan_function_callswould change the matched items, but the complete-history path has no such check. Orphan output pruning inprepare_input_with_session(session_persistence.py490-492) is enabled only withSessionSettings.limit. The existing testtests/memory/test_compaction_model_visibility.pyavoids this case: at line 327 its hook declines compaction at the interruption, with the comment "Avoid compacting the interruption's still-pending tool call.", so no test covers a hook that approves there.Possible fix. Treat a turn that ends in
NextStepInterruptionwith pending calls like a turn with local tool outputs and defer compaction until the resumed turn settles the outputs, or skip automatic compaction whenever the items to compact contain calls without outputs, on the complete-history path as well. A test where the hook approves compaction at the interruption would cover it.