Skip to content

💸 feat: Report Subagent Child-Run Model Usage via Usage Sink - #236

Merged
danny-avila merged 3 commits into
mainfrom
claude/vigilant-snyder-1798fd
Jun 13, 2026
Merged

danny-avila merged 3 commits into
mainfrom
claude/vigilant-snyder-1798fd

Conversation

@danny-avila

Copy link
Copy Markdown
Collaborator

Summary

Subagent child runs spawned by the subagent tool execute via workflow.invoke() with a detached callbacks array — the forwarder only relays custom UI events, so the child's on_chat_model_end events never reach the host's CHAT_MODEL_END handler registered on the parent Run. Hosts (e.g. LibreChat) therefore never see child token usage: those model calls go unbilled and balance enforcement is blind to them. Nested children detach again at every level.

This adds an opt-in usage sink so hosts can observe every model call made inside subagent child runs:

  • Adds SubagentUsageEvent / SubagentUsageSink types: per-call usage tagged with the subagent type, child run ID, child agent ID, parent run ID, model, and provider.
  • Plumbs RunConfig.subagentUsageSink → StandardGraph / MultiAgentGraph → SubagentExecutor (createAgentNode).
  • SubagentExecutor.execute() attaches a usage-capture callback to the child invoke — the child-side equivalent of ModelEndHandler: handleChatModelStart records per-call ls_model_name keyed by callback runId; handleLLMEnd joins it with usage_metadata from the generations and emits through the sink. Captures the agent loop plus auxiliary calls (e.g. child-side summarization).
  • model prefers per-call ls_model_name with fallback to the subagent config's clientOptions model; provider is the config's Providers enum value (not LangSmith's ls_provider string), since hosts key pricing/cache semantics off the enum ('openAI' vs 'openai', 'vertexai' vs 'google_vertexai').
  • Forwards the sink into the child graph's StandardGraphInput, so each nesting level attaches its own capture handler — grandchild usage reports through the same sink (a single top-level handler can never see it, as each invoke replaces the callback chain).
  • Sink dispatch is awaited per model call (awaitHandlers) so all entries are sunk before the parent's tool result resolves; host sink errors are swallowed.
  • Parent-graph calls are NOT reported through the sink — they continue to flow through the registered CHAT_MODEL_END handler, so there is no double-counting.

No behavior change when the sink is omitted.

Testing

  • 6 new SubagentExecutor unit tests: sink forwarding into child graph input, tagged event emission with ls_model_name, configured-model fallback, no-usage skip, no-callback-when-no-sink, sink-error swallowing.
  • New integration spec in src/specs/subagent.test.ts: full Run.processStream with a fake child model that reports usage_metadata on a final stream chunk (mirroring live providers); asserts exactly the child's call reaches the sink, tagged correctly, while the parent's calls stay on the CHAT_MODEL_END path.
  • Full suite green (only pre-existing keyless live-provider specs fail without credentials).
  • Live-verified with real Anthropic and OpenAI credentials via the new src/scripts/subagent-usage-sink.ts script: parent calls landed in collectedUsage, the child's call arrived through the sink with correct model/provider/run tagging and real token counts on both providers.
node -r dotenv/config --loader ./tsconfig-paths-bootstrap.mjs --experimental-specifier-resolution=node ./src/scripts/subagent-usage-sink.ts [--provider anthropic]

Downstream: LibreChat consumes this via createSubagentUsageSink (companion PR) — requires a release > 3.2.33.

@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 34f4d3d18d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +821 to +822
for (const generationGroup of output.generations) {
for (const generation of generationGroup) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Emit one usage event per model call

When a subagent model is configured with multiple completions (for example clientOptions.n > 1 on ChatOpenAI), output.generations[0] contains one ChatGeneration per choice for a single provider request. This nested loop sends the request-level usage_metadata to the sink once per choice, so hosts that bill by summing subagentUsageSink events can overcharge child-run usage. Emit at most one usage event per model callback, or otherwise de-duplicate request-level usage, instead of iterating every candidate generation.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 9786543 — handleLLMEnd now emits at most one event per generation group (break after the first usage-bearing choice), so n > 1 completions no longer multiply the request-level usage. The outer per-prompt loop stays for batched calls, which are genuinely separate requests. Covered by a new unit test (one group × two choices → 1 event; two groups → 2 events).

subagentType,
subagentRunId: childRunId,
subagentAgentId: childAgentId,
parentRunId: this.parentRunId,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve the root run id for nested subagent usage

When allowNested lets a child spawn its own subagent, the child graph constructs a new executor whose parentRunId is the child run id. Passing that value into the usage event means nested usage reaches the host sink with event.runId set to something like <root>_sub_x instead of the original run id, so billing/accounting code that keys the sink events by the host run will drop or misattribute only nested-subagent tokens. The root run id needs to be carried separately through child graphs, or the sink should wrap nested events back to the originating run.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 9786543 — the executor now forwards a wrapped sink into the child graph input that rewrites event.runId to its own parentRunId. The wrappers compose layer-by-layer, so events from any nesting depth surface to the host with the ROOT run id (the emitting child remains identifiable via subagentRunId). SubagentUsageEvent.runId docs updated to state the root-run contract; covered by a unit test simulating a grandchild emission.

Comment on lines +446 to +447
provider: config.agentInputs.provider,
fallbackModel: extractConfiguredModel(config.agentInputs),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Use the provider that actually handled fallback calls

If a subagent's primary model fails and Graph.ts runs a configured fallback provider, this capture handler still tags the usage event with the subagent's original agentInputs.provider. Since the new API explicitly says hosts key pricing/cache semantics from this field, fallback-successful child calls will be billed under the wrong provider (and can also use the wrong fallback model when ls_model_name is absent). Capture the provider/model from the actual callback metadata or from the fallback invocation rather than the static primary config.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 9786543 — rather than trusting static config (or ls_provider, which derived providers like DeepSeek/OpenRouter inherit as 'openai' from their base class), attemptInvoke now stamps INVOKED_PROVIDER (the SDK enum of the provider actually serving the call) onto callback metadata — it's the single funnel for primary, fallback, and summarization invocations. tryFallbackProviders additionally stamps INVOKED_MODEL from the fallback's client options so the model tag doesn't fall back to the primary config when ls_model_name is absent. The capture handler prefers these per-call values with config as last resort. Covered by unit tests on both the stamping (invoke.test.ts) and the capture precedence, plus live-verified.

@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9786543b1b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/tools/subagent/SubagentExecutor.ts Outdated
Comment on lines +878 to +886
sink({
usage,
model,
provider: callProvider,
subagentType,
subagentRunId,
subagentAgentId,
runId: parentRunId,
});

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Await asynchronous usage sink results

When a host implements subagentUsageSink with async billing/persistence, this call drops the returned Promise: handler.awaitHandlers = true only waits for handleLLMEnd itself, which returns immediately, and the surrounding try/catch only catches synchronous throws. In that scenario the parent run can finish before the usage is recorded, and a rejected sink becomes an unhandled rejection instead of being swallowed as the API contract describes.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in a8b495c — SubagentUsageSink now returns void | Promise<void>, handleLLMEnd is async and awaits each dispatch inside the try/catch (so rejections are swallowed per the contract, not unhandled), and the nested-sink wrapper returns the host sink's result so awaits chain through every layer. awaitHandlers then holds the model-call callback until the sink settles — covered by a unit test where an async sink delays + rejects and the dispatch order is asserted.

Comment on lines +843 to +845
const callModel =
asNonEmptyString(metadata?.ls_model_name) ??
asNonEmptyString(metadata?.[Constants.INVOKED_MODEL]);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Use summary model metadata for child usage

When a child run performs summarization with summarizationConfig.model set to a different model, the summarize node passes that value only as metadata.summarization_model, while primary attemptInvoke() only stamps INVOKED_PROVIDER. For providers or custom chat models that do not emit ls_model_name, this code falls through to fallbackModel from the subagent's main clientOptions, so the sink bills the summarization tokens against the wrong model even though the correct model is present in the callback metadata.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in a8b495c — the summarize node now stamps INVOKED_MODEL (alongside the existing summarization_model) onto its invoke metadata when a dedicated summarizer model is configured, using the same attribution mechanism the capture handler already reads; tryFallbackProviders still overrides it per fallback attempt, and self-summarize omits the stamp so the primary-config fallback stays correct. Covered by a node-level test asserting both INVOKED_MODEL and INVOKED_PROVIDER reach the summarizer's model config.

@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. 👍

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@danny-avila
danny-avila merged commit 05f4c3b into main Jun 13, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant