Skip to content

Compaction session compacts a pending approval call, leaving an orphan output #5340

Description

@digit50

Please read this first

Describe the bug

With OpenAIResponsesCompactionSession, automatic compaction runs right after a turn that ends in a tool approval interruption. That turn stored a function_call with no output yet, so the pending call is sent to responses.compact and replaced by the compaction item. After the approval, the resumed run appends the function_call_output, whose function_call is no longer in the session, and every later run sends that unpaired output to the model. prepare_input_with_session only prunes orphan outputs when SessionSettings.limit is set, so on a session without a limit nothing removes it.

Debug information

  • Agents SDK version: 0.23.1 (also checked the code on main at 125efa0, unchanged)
  • Related library versions (optional, e.g. any-llm, litellm, or pydantic): none needed
  • Python version: 3.12.3
  • Operating system: Linux (Pop!_OS 24.04)
  • Model and model provider: none, agents.testing.ScriptedModel and a stub responses.compact client, no network
  • Does the issue reproduce with the latest Agents SDK release? Yes, 0.23.1
  • Does the issue occur consistently or intermittently? Consistently

Output of the script below on 0.23.1:

python: 3.12.3
openai-agents: 0.23.1
compact input at interruption: ['message(user)', 'function_call(c1)']
stored after interruption:    ['compaction']
Skipped compaction because Session history changed after this run appended its items.
stored after resume:          ['compaction', 'function_call_output(c1)', 'message(assistant)']
model input on the next run:  ['compaction', 'function_call_output(c1)', 'message(assistant)', 'message(user)']
expected: compaction is deferred while c1 is pending (or the pair is kept together), so no function_call_output(c1) is stored or sent without function_call(c1)

Repro steps

The compaction client is a stub that returns one compaction item; nothing calls the API.

"""Automatic compaction runs on a turn that ends in an approval interruption, so the pending
function_call is compacted away and the later function_call_output is stored without its call."""

import asyncio
import importlib.metadata
import os
import platform
import tempfile
from types import SimpleNamespace

os.environ.setdefault("OPENAI_API_KEY", "sk-test")
os.environ.setdefault("OPENAI_AGENTS_DISABLE_TRACING", "1")

from agents import Agent, Runner, SQLiteSession, function_tool
from agents.memory import OpenAIResponsesCompactionSession
from agents.testing import ScriptedModel, assistant_message, function_call

print(f"python: {platform.python_version()}")
print(f"openai-agents: {importlib.metadata.version('openai-agents')}")


compact_inputs = []


async def fake_compact(**kwargs):
    compact_inputs.append(shape(kwargs.get("input", [])))
    summary = {"type": "compaction", "id": "cmp-1", "encrypted_content": "summary"}
    return SimpleNamespace(output=[summary], usage=None)


fake_client = SimpleNamespace(responses=SimpleNamespace(compact=fake_compact))


def shape(items):
    out = []
    for item in items:
        item = item if isinstance(item, dict) else item.model_dump(exclude_unset=True)
        kind = item.get("type") or "message"
        if kind == "message":
            out.append(f"message({item.get('role')})")
        elif kind in ("function_call", "function_call_output"):
            out.append(f"{kind}({item['call_id']})")
        else:
            out.append(kind)
    return out


@function_tool(needs_approval=True)
def protected(y: str) -> str:
    """Approval-gated tool."""
    return f"protected:{y}"


async def main() -> None:
    with tempfile.TemporaryDirectory() as tmp:
        store = SQLiteSession("s1", os.path.join(tmp, "history.db"))
        session = OpenAIResponsesCompactionSession(
            "s1",
            store,
            client=fake_client,
            compaction_mode="input",
            should_trigger_compaction=lambda _: True,
        )
        model = ScriptedModel(
            steps=[
                [function_call("protected", {"y": "y"}, call_id="c1")],
                [assistant_message("ok")],
                [assistant_message("second answer")],
            ]
        )
        agent = Agent(name="a", model=model, tools=[protected])

        first = await Runner.run(agent, "hi", session=session)
        print("compact input at interruption:", compact_inputs[-1] if compact_inputs else None)
        print("stored after interruption:   ", shape(await session.get_items()))

        state = first.to_state()
        state.approve(first.interruptions[0])
        await Runner.run(agent, state, session=session)
        print("stored after resume:         ", shape(await session.get_items()))

        await Runner.run(agent, "next question", session=session)
        print("model input on the next run: ", shape(model.calls[-1].input))
        store.close()

    print(
        "expected: compaction is deferred while c1 is pending (or the pair is kept together), "
        "so no function_call_output(c1) is stored or sent without function_call(c1)"
    )


asyncio.run(main())

Expected behavior

A function_call that is waiting for approval should not be compacted away from the output it is about to receive. Either compaction is deferred until the resumed turn has stored the output (as already happens for turns with local tool outputs), or the call and its output stay together, so no function_call_output(c1) is stored or sent without function_call(c1).

Cause. _apply_post_write_compaction (src/agents/run_internal/session_persistence.py 701-714 at 125efa0) defers automatic compaction only when the saved batch has local tool or handoff outputs. A turn that ends in an interruption has pending calls and no outputs, so it falls through to an immediate compaction. In OpenAIResponsesCompactionSession (src/agents/memory/openai_responses_compaction_session.py 397-415) the partial-suffix branch refuses to compact when drop_orphan_function_calls would change the matched items, but the complete-history path has no such check. Orphan output pruning in prepare_input_with_session (session_persistence.py 490-492) is enabled only with SessionSettings.limit. The existing test tests/memory/test_compaction_model_visibility.py avoids this case: at line 327 its hook declines compaction at the interruption, with the comment "Avoid compacting the interruption's still-pending tool call.", so no test covers a hook that approves there.

Possible fix. Treat a turn that ends in NextStepInterruption with pending calls like a turn with local tool outputs and defer compaction until the resumed turn settles the outputs, or skip automatic compaction whenever the items to compact contain calls without outputs, on the complete-history path as well. A test where the hook approves compaction at the interruption would cover it.

Activity

  1. gomission commented on Oct 10, 2026

    @gomission

    Independent extension of the reproduction, with AI assistance: the rejection path also leaves an unpaired output on the next turn, even though the protected tool never executes. I checked both Runner.run and fully consumed Runner.run_streamed after a RunState.to_json() / RunState.from_json() round trip.

    Environment: released openai-agents==0.23.1, Python 3.12.11, openai==3.28.0, Pydantic 2.14.0. Public runner/session APIs, real SQLite, ScriptedModel, and a synthetic responses.compact response; no provider calls.

    Decision Automatic compaction hook Tool dispatches Unpaired pending-c1 output in next-turn model input
    Approve Always true 1 Present in both runner modes
    Reject Always true 0 Present in both runner modes
    Approve Always false 1 Absent in both runner modes
    Reject Always false 0 Absent in both runner modes

    Eight cases executed: four next-turn pairing violations and four controls. In all eight, the immediate resumed model input still has its call/output pairing; the missing call is observable in the persisted history and the following user turn. Rejection remains enforced: this is not a rejected-tool execution.

    For the regression, I would include rejection as well as approval: a denied call still produces a tool output that needs its original call. Deferring only when a backend actually executes would miss that case. The JSON round trip keeps this tied to the public durable state boundary.

    Scope: one process, a stub compaction result and a scripted model. No real compaction API, OpenAI 400 response, separate-worker restore, SDK patch, or full-suite claim. The false compaction hook is a comparison control, not a complete fix.

    Runnable eight-case characterization

    Save as characterize_approval_compaction.py and run:

    uv run --no-project --python 3.12 --with openai-agents==0.23.1 python characterize_approval_compaction.py
    from __future__ import annotations
    
    import asyncio
    import importlib.metadata
    import json
    import platform
    import tempfile
    from pathlib import Path
    from types import SimpleNamespace
    from typing import Any
    
    from agents import Agent, Runner, RunState, SQLiteSession, function_tool
    from agents.memory import OpenAIResponsesCompactionSession
    from agents.testing import ScriptedModel, assistant_message, function_call
    from agents.tracing import set_tracing_disabled
    
    set_tracing_disabled(True)
    
    
    def normalized(items: list[Any]) -> list[dict[str, Any]]:
        return [
            item if isinstance(item, dict) else item.model_dump(exclude_unset=True)
            for item in items
        ]
    
    
    def unpaired_outputs(items: list[Any]) -> list[str]:
        values = normalized(items)
        calls = {item["call_id"] for item in values if item.get("type") == "function_call"}
        return [
            item["call_id"]
            for item in values
            if item.get("type") == "function_call_output" and item["call_id"] not in calls
        ]
    
    
    async def run_case(*, streaming: bool, allow: bool, compact: bool) -> dict[str, Any]:
        dispatches: list[str] = []
        compact_inputs: list[list[dict[str, Any]]] = []
    
        @function_tool(needs_approval=True)
        def protected(value: str) -> str:
            """Return a synthetic value after approval."""
            dispatches.append(value)
            return "synthetic:" + value
    
        async def fake_compact(**kwargs: Any) -> SimpleNamespace:
            compact_inputs.append(normalized(kwargs.get("input", [])))
            return SimpleNamespace(
                output=[
                    {
                        "type": "compaction",
                        "id": "synthetic-compaction",
                        "encrypted_content": "synthetic-summary",
                    }
                ],
                usage=None,
            )
    
        model = ScriptedModel(
            steps=[
                [function_call("protected", {"value": "v"}, call_id="pending-c1")],
                [assistant_message("finished")],
                [assistant_message("next answer")],
            ]
        )
        agent = Agent(name="synthetic-agent", model=model, tools=[protected])
    
        async def run(value: str | RunState, session: Any) -> Any:
            if not streaming:
                return await Runner.run(agent, value, session=session)
            result = Runner.run_streamed(agent, value, session=session)
            async for _ in result.stream_events():
                pass
            return result
    
        with tempfile.TemporaryDirectory(prefix="synthetic-compaction-") as directory:
            store = SQLiteSession("synthetic-session", str(Path(directory) / "history.db"))
            try:
                session = OpenAIResponsesCompactionSession(
                    "synthetic-session",
                    store,
                    client=SimpleNamespace(responses=SimpleNamespace(compact=fake_compact)),
                    compaction_mode="input",
                    should_trigger_compaction=lambda _: compact,
                )
                first = await run("first question", session)
                assert len(first.interruptions) == 1 and not dispatches
                stored_at_pause = normalized(await session.get_items())
                state = first.to_state()
                state_json = state.to_json()
                restored = await RunState.from_json(agent, state_json)
                interruption = first.interruptions[0]
                if allow:
                    restored.approve(interruption)
                else:
                    restored.reject(interruption)
                await run(restored, session)
                stored_after_resume = normalized(await session.get_items())
                resume_model_input = normalized(model.calls[-1].input)
                await run("next question", session)
                next_model_input = normalized(model.calls[-1].input)
                assert len(dispatches) == (1 if allow else 0)
                return {
                    "streaming": streaming,
                    "decision": "approve" if allow else "reject",
                    "compaction_enabled": compact,
                    "runstate_json_roundtrip": True,
                    "dispatches": len(dispatches),
                    "compact_invocations": len(compact_inputs),
                    "stored_at_pause": stored_at_pause,
                    "stored_after_resume": stored_after_resume,
                    "resume_model_input": resume_model_input,
                    "next_model_input": next_model_input,
                    "resume_unpaired_outputs": unpaired_outputs(resume_model_input),
                    "next_turn_unpaired_outputs": unpaired_outputs(next_model_input),
                }
            finally:
                store.close()
    
    
    async def main() -> None:
        cases = []
        for streaming in (False, True):
            for allow in (True, False):
                for compact in (True, False):
                    cases.append(
                        await run_case(streaming=streaming, allow=allow, compact=compact)
                    )
        violations = [case for case in cases if case["next_turn_unpaired_outputs"]]
        report = {
            "python": platform.python_version(),
            "openai_agents": importlib.metadata.version("openai-agents"),
            "openai": importlib.metadata.version("openai"),
            "pydantic": importlib.metadata.version("pydantic"),
            "cases_executed": len(cases),
            "cases_with_next_turn_unpaired_outputs": len(violations),
            "cases": cases,
            "scope": (
                "Public SDK Runner, real SQLite session and RunState JSON round trip; "
                "ScriptedModel and synthetic compaction response, one process. "
                "No model/compaction API calls, separate-process restore, SDK fix or full suite."
            ),
        }
        Path(__file__).with_name("verification.json").write_text(
            json.dumps(report, indent=2) + "\n"
        )
        print(
            json.dumps(
                {
                    **{key: value for key, value in report.items() if key != "cases"},
                    "case_summaries": [
                        {
                            key: case[key]
                            for key in (
                                "streaming",
                                "decision",
                                "compaction_enabled",
                                "dispatches",
                                "resume_unpaired_outputs",
                                "next_turn_unpaired_outputs",
                            )
                        }
                        for case in cases
                    ],
                },
                indent=2,
            )
        )
    
    
    if __name__ == "__main__":
        asyncio.run(main())
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions