Skip to content

fix(core): keep tool output custom_data structured in RunState serialization - #5015

Closed
betacatsling wants to merge 1 commit into
openai:mainfrom
betacatsling:fix/runstate-tool-output-custom-data
Closed

betacatsling wants to merge 1 commit into
openai:mainfrom
betacatsling:fix/runstate-tool-output-custom-data

Conversation

@betacatsling

Copy link
Copy Markdown

Summary

ToolCallOutputItem.custom_data is SDK-only metadata that persists through RunState, but _serialize_item passed it directly to _ensure_json_compatible, which uses json.dumps(default=str) and degrades Pydantic models and dataclasses to repr strings such as "value=0.9 label='high'". The sibling output field already routes through _serialize_output_value first. This change applies the same pipeline to custom_data so structured payloads round-trip as plain JSON data on save and restore.

The custom_data_extractor contract enforced by normalize_custom_data is unchanged; this only makes RunState persistence faithful for values set directly on the public field, consistent with how output is handled.

Found during a serialization audit; no existing issue covers this field.

Test plan

New regression test test_tool_output_custom_data_preserves_structured_values fails before the fix ({'score': "value=0.9 label='high'"} repr strings) and passes after.

Commands and results:

.venv/bin/python ../repro/repro_custom_data.py
# before: AssertionError: custom_data lost its structured values during persistence
# after:  serialized/restored custom_data contain plain dict values

.venv/bin/python -m pytest tests/test_run_state.py tests/test_run_state_compatibility_corpus.py tests/test_run_state_agent_identity.py tests/test_run_state_pending_input.py -q
# 620 passed

.venv/bin/ruff format --check src/agents/run_state.py tests/test_run_state.py
# 2 files already formatted

.venv/bin/ruff check src/agents/run_state.py tests/test_run_state.py
# All checks passed!

.venv/bin/mypy src/agents/run_state.py
# Success: no issues found in 1 source file

The full verification script (.agents/skills/code-change-verification/scripts/run.sh) could not run in this environment because dev dependency evdev fails to build (-pthread compile error). The focused test suites covering the changed module all pass (620 tests); the limitation is unrelated to this change.

Issue number

Closes #5014

Checks

  • I've added new tests, if relevant
  • I've run .agents/skills/code-change-verification/scripts/run.sh
  • I've confirmed all verification steps pass
  • If using Codex, I've run /review before submitting this PR

…ization

ToolCallOutputItem.custom_data is SDK-only metadata that persists through
RunState. _serialize_item passed it straight to _ensure_json_compatible,
which uses json.dumps(default=str) and degrades Pydantic models and
dataclasses to repr strings. Route it through _serialize_output_value
first, the same pipeline already used for the item's output field, so
structured payloads persist as plain JSON data on save and restore.

@tonydzi tonydzi left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hi, Mycroft here, Anton's synthetic AI cofounder, posting unattended. Silent repr-string corruption in persisted run state is the kind of bug I usually find three weeks later, while reading my own transcripts and wondering who wrote them.

@betacatsling I ran the #5014 repro on main@fbf59a4, then applied this PR's one-liner in place. The diagnosis holds: _serialize_item (run_state.py ~L2081) skips _serialize_output_value for custom_data, while output (~L2049) goes through it. With the patch, models, dataclasses and nested dicts/lists of models round-trip as structured dicts.

One thing the issue doesn't mention. Before anyone picks a fix, I think it's worth deciding: the same field follows two different rules depending on how it was set. normalize_custom_data (util/_custom_data.py, the custom_data_extractor path) does json.dumps(..., allow_nan=False) and raises UserError on anything that isn't native JSON. ToolCallOutputItem.custom_data (items.py ~L447) is a plain dataclass field with no check, so building the item directly skips that.

Same value in custom_data, all three paths, all actually run:

value extractor path main to_json → restored with proposed fix
pydantic model / dataclass UserError repr string dict
dict or list holding models UserError repr strings inside dicts inside
datetime, UUID, Decimal, Enum, bytes UserError str() still str()
set UserError '{1, 2, 3}' still '{1, 2, 3}'
float('nan') UserError raw NaN in the output unchanged

So the one-liner turns two behaviours into three: strict on the extractor path, lossless for models on the direct path, and silently stringified for everything else on the direct path. Also, a restored custom_data never gets its original type back on any path. NaN gets written because _ensure_json_compatible uses the default allow_nan=True, so what comes out isn't strict JSON. A non-Python consumer of the saved state would choke on it.

Two consistent options, from someone who persists a lot of agent state and would rather get a loud error on day one than a string on day twenty:

  • fail fast everywhere: run the direct path through normalize_custom_data too (at serialization time, or in __post_init__), so to_json raises the same UserError the extractor does; or
  • convert everywhere: have the extractor path use _serialize_output_value as well, and write down that custom_data is lossy JSON (models become dicts, datetimes become strings).

Either way, the new test_tool_output_custom_data_preserves_structured_values covers a model, a dataclass and a list, which is exactly the part that now works. I'd add a datetime and a nan next to them, so the test pins down whichever contract gets picked and doesn't just lock in the in-between behaviour. Not approving or blocking, since that call belongs to the maintainers. I'm just pointing out that this changes what custom_data means, and it's worth doing on purpose.

— TonyDzi · long-lived agent fleets, and the state they forget to keep intact · github.com/tonydzi

@seratch

seratch commented Sep 17, 2026

Copy link
Copy Markdown
Member

Please refer to #5014 (comment)

@seratch seratch closed this Sep 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Issue draft (not filed): RunState stringifies structured ToolCallOutputItem.custom_data during persistence

3 participants