Skip to content

Desktop/app-server history replay omits persisted commandExecution items #28162

Description

@VincentAdamNemessisX

Summary

Codex Desktop/app-server history replay appears to omit persisted command execution items after a session is reloaded. The raw rollout JSONL still contains the command tool calls and outputs, but thread/read with includeTurns: true returns no commandExecution items for the same thread.

Environment

  • Codex CLI: codex-cli 0.133.0
  • Codex Desktop / app-server protocol: thread/read with includeTurns: true
  • Platform: macOS

Evidence

  • Raw rollout JSONL contains persisted command execution data, including response_item payloads with type: "function_call" and name: "exec_command", matching function_call_output payloads, and in some sessions event_msg payloads with type: "exec_command_end".
  • Across a local rollout corpus, the raw files include 47,534 exec_command function calls, 56,291 function call outputs, and 5,603 exec_command_end events.
  • For one sampled thread, raw JSONL contained thousands of tool/function rows. The SQLite thread index pointed to the correct rollout path. However thread/read includeTurns=true returned 130 turns with item counts such as fileChange: 250, mcpToolCall: 53, agentMessage: 757, and commandExecution: 0.
  • Additional sampled sessions with only legacy response_item.function_call(name=exec_command) rows and sessions with newer event_msg.exec_command_end rows also replayed with commandExecution: 0.
  • This makes command execution details disappear from the Desktop UI after reload/switch even though the underlying rollout JSONL is still intact.

Expected behavior

History replay should reconstruct command execution UI items, or an equivalent expandable tool-call item, from persisted rollout rows such as:

{"type":"response_item","payload":{"type":"function_call","name":"exec_command","call_id":"...","arguments":"{...}"}}
{"type":"response_item","payload":{"type":"function_call_output","call_id":"...","output":"..."}}
{"type":"event_msg","payload":{"type":"exec_command_end","call_id":"...","command":["/bin/zsh","-lc","..."],"cwd":"...","exit_code":0,"status":"completed"}}

Actual behavior

After reload/history replay, command execution details are missing from the returned thread items and from the Desktop UI. Other history items such as assistant messages, file changes, and some MCP/tool items can still appear.

Reproduction outline

  1. Create or identify a Codex session that has shell command executions.
  2. Confirm the raw rollout JSONL contains exec_command function calls and matching outputs, and/or exec_command_end events.
  3. Start app-server and call thread/read with includeTurns: true for that thread.
  4. Count returned turn items. The raw rollout has command execution data, but the replayed turns contain zero commandExecution items.

Notes

  • This report intentionally avoids including local paths, credentials, command output, or private session content.
  • The problem looks like a history replay/normalization issue rather than physical JSONL loss: the raw rollout files remain intact and the SQLite threads.rollout_path values point to the correct files.

Activity

  1. added
    bugSomething isn't working
    appIssues related to the Codex desktop app
    app-serverIssues involving app server protocol or interfaces
    sessionIssues involving session (thread) management, resuming, forking, naming, archiving
    on Jun 14, 2026
  2. wgu9 commented on Jun 15, 2026

    @wgu9

    I took a look at the current replay reducer. The missing-command-history behavior seems to split into two cases.

    For newer rollout rows that contain event_msg values, current main already has a replay path:

    • thread/read includeTurns loads rollout history and calls build_api_turns_from_rollout_items(...)
    • ThreadHistoryBuilder::handle_event(...) handles EventMsg::ExecCommandBegin and EventMsg::ExecCommandEnd
    • handle_exec_command_end(...) calls build_command_execution_end_item(...), which maps the core event into ThreadItem::CommandExecution with aggregated_output, exit_code, and duration_ms
    • there are existing tests in app-server-protocol/src/protocol/thread_history.rs asserting that an ExecCommandEndEvent replays into a completed CommandExecution item

    The gap I do see is the legacy-only shape mentioned in the report: rollout rows that have only response_item.function_call(name = "exec_command") plus matching response_item.function_call_output, without the newer EventMsg::ExecCommandBegin/End records. In current ThreadHistoryBuilder::handle_response_item(...), response-item replay only handles user Message rows that parse as hook prompts. Non-message response items, including function_call and function_call_output, return early and do not contribute any ThreadItem.

    That explains why older sessions with only persisted Responses API tool-call rows can replay with commandExecution: 0 even though the raw JSONL still contains command calls and outputs.

    A minimal fix path would probably be:

    1. Keep the existing EventMsg::ExecCommand* replay as the canonical path for newer rollouts.
    2. Add a legacy fallback in ThreadHistoryBuilder that pairs ResponseItem::FunctionCall { name: "exec_command", call_id, arguments, ... } with ResponseItem::FunctionCallOutput { call_id, output }.
    3. Reconstruct a best-effort ThreadItem::CommandExecution from those paired rows, using the parsed command/cwd from the arguments when available and the function-call output as aggregated_output.
    4. Avoid duplicate items if both legacy response rows and newer ExecCommandEnd event rows exist for the same call_id; the event row should win because it carries richer status, cwd, parsed command actions, exit code, and duration.

    A focused regression test could build rollout items containing only the legacy response_item pair for an exec_command call and assert build_turns_from_rollout_items(...) returns one ThreadItem::CommandExecution. A second test should include both the legacy pair and an ExecCommandEndEvent for the same call_id and assert only the event-derived command item survives.

    Per docs/contributing.md, I am not opening an unsolicited PR, but this looks like a contained replay-normalization fix if maintainers want it.

  3. VincentAdamNemessisX commented on Jun 17, 2026

    @VincentAdamNemessisX
    Author

    I took a look at the current replay reducer. The missing-command-history behavior seems to split into two cases.

    For newer rollout rows that contain event_msg values, current main already has a replay path:

    • thread/read includeTurns loads rollout history and calls build_api_turns_from_rollout_items(...)
    • ThreadHistoryBuilder::handle_event(...) handles EventMsg::ExecCommandBegin and EventMsg::ExecCommandEnd
    • handle_exec_command_end(...) calls build_command_execution_end_item(...), which maps the core event into ThreadItem::CommandExecution with aggregated_output, exit_code, and duration_ms
    • there are existing tests in app-server-protocol/src/protocol/thread_history.rs asserting that an ExecCommandEndEvent replays into a completed CommandExecution item

    The gap I do see is the legacy-only shape mentioned in the report: rollout rows that have only response_item.function_call(name = "exec_command") plus matching response_item.function_call_output, without the newer EventMsg::ExecCommandBegin/End records. In current ThreadHistoryBuilder::handle_response_item(...), response-item replay only handles user Message rows that parse as hook prompts. Non-message response items, including function_call and function_call_output, return early and do not contribute any ThreadItem.

    That explains why older sessions with only persisted Responses API tool-call rows can replay with commandExecution: 0 even though the raw JSONL still contains command calls and outputs.

    A minimal fix path would probably be:

    1. Keep the existing EventMsg::ExecCommand* replay as the canonical path for newer rollouts.
    2. Add a legacy fallback in ThreadHistoryBuilder that pairs ResponseItem::FunctionCall { name: "exec_command", call_id, arguments, ... } with ResponseItem::FunctionCallOutput { call_id, output }.
    3. Reconstruct a best-effort ThreadItem::CommandExecution from those paired rows, using the parsed command/cwd from the arguments when available and the function-call output as aggregated_output.
    4. Avoid duplicate items if both legacy response rows and newer ExecCommandEnd event rows exist for the same call_id; the event row should win because it carries richer status, cwd, parsed command actions, exit code, and duration.

    A focused regression test could build rollout items containing only the legacy response_item pair for an exec_command call and assert build_turns_from_rollout_items(...) returns one ThreadItem::CommandExecution. A second test should include both the legacy pair and an ExecCommandEndEvent for the same call_id and assert only the event-derived command item survives.

    Per docs/contributing.md, I am not opening an unsolicited PR, but this looks like a contained replay-normalization fix if maintainers want it.

    Okay, thanks for your reply.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    appIssues related to the Codex desktop appapp-serverIssues involving app server protocol or interfacesbugSomething isn't workingsessionIssues involving session (thread) management, resuming, forking, naming, archiving

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions