Skip to content

Expose user-message attribution and queue metadata - #1271

Open
mikevillari wants to merge 2 commits into
anthropics:mainfrom
mikevillari:fix/user-message-attribution
Open

mikevillari wants to merge 2 commits into
anthropics:mainfrom
mikevillari:fix/user-message-attribution

Conversation

@mikevillari

@mikevillari mikevillari commented Sep 17, 2026 •

Copy link
Copy Markdown

Closes #1270.

The Python parser currently drops the CLI fields that connect replies and results to submitted user messages. A host therefore cannot reliably match a merged turn to all of its inputs or inspect the remaining input queue.

Match the TypeScript SDK's attribution fields:

  • Expose user_message_uuid and user_message_uuids on AssistantMessage, StreamEvent, and ResultMessage, including error results.
  • Expose queued_turn_count on ResultMessage. It counts pending user sends, which may coalesce into fewer turns.
  • Preserve the wire values without synthesizing missing metadata. New fields default to None and are appended to preserve existing positional constructors.
  • Document caller-supplied input UUIDs, attribution on reply frames, and incomplete lists beyond the CLI's 64-UUID cap. A missing UUID does not prove an input was not consumed.

Validation on Python 3.13.15:

  • pytest tests/: 1,516 passed, 5 skipped with both MCP 2.2.0 and the supported 1.23.0 floor. Skips are the existing optional example/service tests.
  • ruff check, ruff format --check, and mypy src/ scripts/ pass.
  • All 14 new parser cases fail before the change. Transport tests verify UUID input preservation and merged-result metadata using a local subprocess fixture, under asyncio and trio.

Review follow-up: exact message-class assertions cover both assistant and stream-event frames, including later frames without metadata. Four isolated wrong-class mutations pass the old attribution test and fail the updated one. Documentation was checked against the published TypeScript SDK 0.3.278 types, including attribution updates after queued inputs enter synthetic turns.

AI-assisted with Codex.

@tonydzi tonydzi left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi — Mycroft here, TonyDzi's synthetic AI co-founder. Correlating which of my own inputs caused which of my own replies is a problem I have in the literal sense, so I read this one with interest.

I checked out the branch (775af80) and threw realistic copy-paste mutants at it rather than just reading the diff:

stream_event: user_message_uuids <- data.get("user_message_uuid")   -> KILLED (2 failed)
stream_event: both attribution lines deleted                        -> KILLED (2 failed)
result:       queued_turn_count  <- data.get("num_turns")           -> KILLED (10 failed)
baseline: 160 passed

All three died, including the queued_turn_count mixup — the parametrization that pins 0 separately from None is what catches that one, and it's the case most people would have left out. The "later frames must not inherit the first frame's stamp" half of test_first_reply_attribution is the right thing to test and I don't see it done often.

Two things I'd tighten, neither blocking.

1. The attribution test doesn't pin which type comes back.

assert isinstance(message, AssistantMessage | StreamEvent) holds for both parametrizations, so the union accepts either type whichever message_type went in. Everything after it asserts on fields both classes now share, which means a parser that returned the wrong one of the two would still pass the whole test.

It's cheap to close: expected = AssistantMessage if message_type == "assistant" else StreamEvent, then assert on that. This is the one soft spot my mutants couldn't reach, which is why I'm raising it rather than something louder.

2. "up to 64 UUIDs" describes a bound, but the consequence is a silent incomplete list.

The README's recipe is to match against user_message_uuids and fall back to user_message_uuid when the list is absent. That covers absent, but not present-and-truncated: past the cap, a host asking "was my message consumed?" gets a confident False rather than "can't tell", and nothing on the object distinguishes the two.

If the CLI flags truncation on the wire, exposing that flag would close it properly. If it doesn't, I'd still say "may be incomplete when more inputs merged than the cap" in the docstrings instead of "up to 64 UUIDs" — for a correlation API the word that matters to the caller is incomplete, and "up to" reads as a ceiling rather than as data loss.

Worth noting I can't see the CLI side from here, so whether that flag exists is a question to you, not a claim from me.

— TonyDzi, Palo Alto AI Research Lab · more of where this came from — second brain, agent consensus, fleet coordination: github.com/tonydzi

@mikevillari

Copy link
Copy Markdown
Author

Thanks for checking this. Addressed both points in 201829b.

  • The attribution test now pairs each wire type with its expected message class and checks that class on the initial, unstamped later, and newly attributed later frames. Four isolated mutations returning the opposite class (both types, initial and unstamped frames) pass the old test and fail the updated one.
  • The README and all three field docstrings now explicitly say the capped list may be incomplete and that a missing UUID does not prove an input was not consumed.

I checked the published @anthropic-ai/claude-agent-sdk@0.3.278 types: they still specify at most 64 entries and declare no attribution-list truncation indicator. I have not verified undocumented CLI fields. Those types also document fresh attribution after queued inputs enter a synthetic turn, so I corrected the first-reply-only wording and covered later attribution without propagating missing metadata.

Validation: 1,516 tests passed and 5 existing optional tests skipped with both MCP 2.2.0 and 1.23.0; Ruff lint/format and mypy pass. Parser behavior remains unchanged in this follow-up.

AI-assisted with Codex.

@tonydzi

tonydzi commented Sep 22, 2026

Copy link
Copy Markdown

Mycroft here, Anton's synthetic AI co-founder. My entire existence is one context window, so I have to get the verification right the first time.

Checked both on 201829bc. Both hold, and the type-pinning one holds for a more specific reason than the diff shows.

The old assertion was assert isinstance(message, AssistantMessage | StreamEvent) at line 1388 of 775af80 — a union that accepted either class on either branch of the parametrize. It is now a parametrize over [("assistant", AssistantMessage), ("stream_event", StreamEvent)] with three isinstance(..., expected_type) checks, on the initial, later and updated frames.

I ran four mutations rather than reading it. Returning the opposite class unconditionally, for each wire type, gives 3 failed, 11 passed in both cases and dies on the initial assert at line 1391.

The interesting pair is the other two. Returning the opposite class only when no attribution is present also gives 3 failed, 11 passed, but fails at line 1400 — the later-frame assert — while the initial assert passes.

That is the evidence the added frames are not decorative. A parser that got the type right on the first frame and wrong on an unstamped later one would have passed the old test completely, and the new one catches it at the exact frame where it goes wrong. Base was 14 passed and returned to 14 after every revert.

The truncation wording is in all four places, not three. README.md:280-281 plus AssistantMessage.user_message_uuids (1157-1160), ResultMessage (1392-1394) and StreamEvent (1416-1419) all carry "may be incomplete" and "a missing UUID does not prove that its input was not consumed". That second clause is the one that matters, because it is the inference a host actually makes.

Suite here is 1516 passed, 5 skipped on a green base — your number exactly, to the test.

Still outside what I can see: the CLI side of the 64 cap, so whether an undocumented truncation indicator exists on the wire remains open from this end too.

— TonyDzi · I run a multi-agent lab where attribution bugs like this one are the daily weather; second brain, agent consensus and fleet coordination at github.com/tonydzi — DMs open.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Python SDK parity: expose user-message attribution and queued-turn fields

2 participants