Repository navigation
💸 feat: Report Subagent Child-Run Model Usage via Usage Sink - #236
Conversation
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 34f4d3d18d
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| for (const generationGroup of output.generations) { | ||
| for (const generation of generationGroup) { |
There was a problem hiding this comment.
Emit one usage event per model call
When a subagent model is configured with multiple completions (for example clientOptions.n > 1 on ChatOpenAI), output.generations[0] contains one ChatGeneration per choice for a single provider request. This nested loop sends the request-level usage_metadata to the sink once per choice, so hosts that bill by summing subagentUsageSink events can overcharge child-run usage. Emit at most one usage event per model callback, or otherwise de-duplicate request-level usage, instead of iterating every candidate generation.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 9786543 — handleLLMEnd now emits at most one event per generation group (break after the first usage-bearing choice), so n > 1 completions no longer multiply the request-level usage. The outer per-prompt loop stays for batched calls, which are genuinely separate requests. Covered by a new unit test (one group × two choices → 1 event; two groups → 2 events).
| subagentType, | ||
| subagentRunId: childRunId, | ||
| subagentAgentId: childAgentId, | ||
| parentRunId: this.parentRunId, |
There was a problem hiding this comment.
Preserve the root run id for nested subagent usage
When allowNested lets a child spawn its own subagent, the child graph constructs a new executor whose parentRunId is the child run id. Passing that value into the usage event means nested usage reaches the host sink with event.runId set to something like <root>_sub_x instead of the original run id, so billing/accounting code that keys the sink events by the host run will drop or misattribute only nested-subagent tokens. The root run id needs to be carried separately through child graphs, or the sink should wrap nested events back to the originating run.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 9786543 — the executor now forwards a wrapped sink into the child graph input that rewrites event.runId to its own parentRunId. The wrappers compose layer-by-layer, so events from any nesting depth surface to the host with the ROOT run id (the emitting child remains identifiable via subagentRunId). SubagentUsageEvent.runId docs updated to state the root-run contract; covered by a unit test simulating a grandchild emission.
| provider: config.agentInputs.provider, | ||
| fallbackModel: extractConfiguredModel(config.agentInputs), |
There was a problem hiding this comment.
Use the provider that actually handled fallback calls
If a subagent's primary model fails and Graph.ts runs a configured fallback provider, this capture handler still tags the usage event with the subagent's original agentInputs.provider. Since the new API explicitly says hosts key pricing/cache semantics from this field, fallback-successful child calls will be billed under the wrong provider (and can also use the wrong fallback model when ls_model_name is absent). Capture the provider/model from the actual callback metadata or from the fallback invocation rather than the static primary config.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 9786543 — rather than trusting static config (or ls_provider, which derived providers like DeepSeek/OpenRouter inherit as 'openai' from their base class), attemptInvoke now stamps INVOKED_PROVIDER (the SDK enum of the provider actually serving the call) onto callback metadata — it's the single funnel for primary, fallback, and summarization invocations. tryFallbackProviders additionally stamps INVOKED_MODEL from the fallback's client options so the model tag doesn't fall back to the primary config when ls_model_name is absent. The capture handler prefers these per-call values with config as last resort. Covered by unit tests on both the stamping (invoke.test.ts) and the capture precedence, plus live-verified.
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9786543b1b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| sink({ | ||
| usage, | ||
| model, | ||
| provider: callProvider, | ||
| subagentType, | ||
| subagentRunId, | ||
| subagentAgentId, | ||
| runId: parentRunId, | ||
| }); |
There was a problem hiding this comment.
Await asynchronous usage sink results
When a host implements subagentUsageSink with async billing/persistence, this call drops the returned Promise: handler.awaitHandlers = true only waits for handleLLMEnd itself, which returns immediately, and the surrounding try/catch only catches synchronous throws. In that scenario the parent run can finish before the usage is recorded, and a rejected sink becomes an unhandled rejection instead of being swallowed as the API contract describes.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in a8b495c — SubagentUsageSink now returns void | Promise<void>, handleLLMEnd is async and awaits each dispatch inside the try/catch (so rejections are swallowed per the contract, not unhandled), and the nested-sink wrapper returns the host sink's result so awaits chain through every layer. awaitHandlers then holds the model-call callback until the sink settles — covered by a unit test where an async sink delays + rejects and the dispatch order is asserted.
| const callModel = | ||
| asNonEmptyString(metadata?.ls_model_name) ?? | ||
| asNonEmptyString(metadata?.[Constants.INVOKED_MODEL]); |
There was a problem hiding this comment.
Use summary model metadata for child usage
When a child run performs summarization with summarizationConfig.model set to a different model, the summarize node passes that value only as metadata.summarization_model, while primary attemptInvoke() only stamps INVOKED_PROVIDER. For providers or custom chat models that do not emit ls_model_name, this code falls through to fallbackModel from the subagent's main clientOptions, so the sink bills the summarization tokens against the wrong model even though the correct model is present in the callback metadata.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in a8b495c — the summarize node now stamps INVOKED_MODEL (alongside the existing summarization_model) onto its invoke metadata when a dedicated summarizer model is configured, using the same attribution mechanism the capture handler already reads; tryFallbackProviders still overrides it per fallback attempt, and self-summarize omits the stamp so the primary-config fallback stays correct. Covered by a node-level test asserting both INVOKED_MODEL and INVOKED_PROVIDER reach the summarizer's model config.
|
@codex review |
|
Codex Review: Didn't find any major issues. 👍 ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
Summary
Subagent child runs spawned by the
subagenttool execute viaworkflow.invoke()with a detached callbacks array — the forwarder only relays custom UI events, so the child'son_chat_model_endevents never reach the host'sCHAT_MODEL_ENDhandler registered on the parentRun. Hosts (e.g. LibreChat) therefore never see child token usage: those model calls go unbilled and balance enforcement is blind to them. Nested children detach again at every level.This adds an opt-in usage sink so hosts can observe every model call made inside subagent child runs:
SubagentUsageEvent/SubagentUsageSinktypes: per-call usage tagged with the subagenttype, child run ID, child agent ID, parent run ID, model, and provider.RunConfig.subagentUsageSink→StandardGraph/MultiAgentGraph→SubagentExecutor(createAgentNode).SubagentExecutor.execute()attaches a usage-capture callback to the child invoke — the child-side equivalent ofModelEndHandler:handleChatModelStartrecords per-callls_model_namekeyed by callback runId;handleLLMEndjoins it withusage_metadatafrom the generations and emits through the sink. Captures the agent loop plus auxiliary calls (e.g. child-side summarization).modelprefers per-callls_model_namewith fallback to the subagent config'sclientOptionsmodel;provideris the config'sProvidersenum value (not LangSmith'sls_providerstring), since hosts key pricing/cache semantics off the enum ('openAI'vs'openai','vertexai'vs'google_vertexai').StandardGraphInput, so each nesting level attaches its own capture handler — grandchild usage reports through the same sink (a single top-level handler can never see it, as eachinvokereplaces the callback chain).awaitHandlers) so all entries are sunk before the parent's tool result resolves; host sink errors are swallowed.CHAT_MODEL_ENDhandler, so there is no double-counting.No behavior change when the sink is omitted.
Testing
SubagentExecutorunit tests: sink forwarding into child graph input, tagged event emission withls_model_name, configured-model fallback, no-usage skip, no-callback-when-no-sink, sink-error swallowing.src/specs/subagent.test.ts: fullRun.processStreamwith a fake child model that reportsusage_metadataon a final stream chunk (mirroring live providers); asserts exactly the child's call reaches the sink, tagged correctly, while the parent's calls stay on theCHAT_MODEL_ENDpath.src/scripts/subagent-usage-sink.tsscript: parent calls landed incollectedUsage, the child's call arrived through the sink with correct model/provider/run tagging and real token counts on both providers.Downstream: LibreChat consumes this via
createSubagentUsageSink(companion PR) — requires a release > 3.2.33.