Repository navigation
Conversation
When a model call or tool failed, OpenTelemetry marked every span the exception passed through as ERROR, but only execute_tool said which error it was. The invocation, invoke_workflow, invoke_agent, call_llm and generate_content spans had no error.type, even though the metrics for the same operations record it. Open these spans through a small helper that sets error.type from resolve_error_type() when an exception escapes, and set it where a workflow node hands its failure back as data. The functional goldens for the four error cases are re-recorded, and the generate_content divergence from the OTel instrumentor is now a value difference (429 vs ClientError), the same one already accepted for the duration metric. Fixes google#7494
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Link to Issue or Description of Change
1. Link to an existing issue (if applicable):
Problem:
When a model call or tool fails, OpenTelemetry marks every span the exception passes through as ERROR, but only
execute_toolsetserror.type. Theinvocation/invoke_workflow,invoke_agent,call_llmandgenerate_contentspans end in ERROR without saying which error it was, even though the metrics for the same operations record it (429,ValueError, ...).test_error_status_implies_error_typetracked this as a strict xfail for four cases.Solution:
tracing.start_as_current_span(), a small wrapper aroundtracer.start_as_current_span()that setserror.typefromresolve_error_type()when an exception escapes the span, then re-raises. It only catchesException, the same cases where OpenTelemetry sets the ERROR status.invocation,invoke_agent,call_llm,generate_content,invoke_workflowandinvoke_nodespans are now opened through it.node_tracing.py, where a node hands its failure back as data and the workflow span is marked ERROR explicitly,error.typeis set as well.tracingandnode_tracing, which is the same object.node_tracingno longer importstracer, so the harness now patchestracing.traceronce.python -m tests.unittests.telemetry.regenerate.generate_contentdivergence forerror.typewas anadk_bug("leaveserror.typeoff the inference span"). ADK now reports429where the OTel instrumentor reportsClientError, so I changed it todesired_behaviorwith the same reasoning already used for the duration metric'serror.type.Failed spans gain one attribute; nothing is renamed or removed.
Testing Plan
Unit Tests:
test_error_status_implies_error_typenow runs as a normal test for every case and passes.The
error.typeentry in_FACTS_MISSING_FROM_SPANSstays as an xfail. What's left there are the skill-script cases: the failing scripts' spans already carryerror.type, but the test pairs those metric points with the span of the script that succeeded, since the metric doesn't record which script failed. That's separate from this change.Manual End-to-End (E2E) Tests:
Covered by the functional telemetry tests, which run real agent invocations against an in-memory exporter. The diff in the four re-recorded goldens shows the new
error.typeattribute on each span in the failing chain.Checklist