Repository navigation
Conversation
2b9dc58 to
27768bf
Compare
Scripts for single and fragmented blocks, multi-block requests, differing source/target dialect and requested/served model, and section rejection with param. Expected to fail until llm.idl and LlmFunctions gain the new fields (#2633). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K1XDb69ajDTbvbSyUfPGFq
LlmDataEx gains message plus reserved attributes and extension slots and drops logProbability; LlmBeginEx gains a dialect extension slot; LlmError gains param. LlmFunctions builders and matchers cover the new fields. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K1XDb69ajDTbvbSyUfPGFq
…nical.transformed Matches the .transformed naming used for dialect transformation scenarios. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K1XDb69ajDTbvbSyUfPGFq
…ulti.block Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01K1XDb69ajDTbvbSyUfPGFq
…request stream Scripts and ITs for decoding an OpenAI Chat Completions request into typed blocks (llm server) and encoding blocks into one (llm client), written before any runtime change. Sections covered: system-instruction (system and developer roles), user-text, user-image, user-document, user-audio, assistant-text, assistant-refusal, tool-definition, tool-call, tool-result, mixed-content messages and a full conversation. Text sections carry the text as payload; part and tool objects carry the verbatim OpenAI JSON object; a tool-result carries its tool_call_id as the block extension. Also covers member order independence (role, model and tool_call_id after content), a rejection when content before role exceeds the decode slot, fragmented 10k and 100k blocks, and an openai to openai round trip. Spec peer-to-peer ITs pass. Runtime ITs are expected to fail until the section decode and encode are implemented. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
Frees the dialect type names for a per-dialect handler SPI. The existing dialect interface, context, factory SPI, resolver, event and terminator types, and the openai and anthropic implementations and their test fixtures, become LlmLegacy*. Service registrations follow the renamed SPI. No behavior change. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
Adds a dialect SPI so each dialect can own its stream handling: LlmDialectFactorySpi creates an LlmDialect once per engine, which supplies an LlmDialectContext per worker, which attaches to a binding and returns an LlmDialectHandler that detects its dialect from the request headers and opens streams. LlmServerFactory and LlmClientFactory become lean dispatchers. The server attaches every dialect that supports the binding and selects the configured dialect, or the single handler that detects the request, rejecting when none or several do. The client attaches only the configured dialect. The previous stream implementations move behind legacy handlers, one per dialect that is not registered through the new SPI, so openai and anthropic behave exactly as before and every existing IT is unchanged. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
LlmDialectContext.attach takes the engine's BindingConfig instead of the internal LlmBindingConfig, so the exported dialect package no longer references an internal type. Each dialect builds what it needs from the binding; the lean server and client factories only read the dialect option. The json and agrona modules are required transitively, since the exported handler signature exposes their types. No behavior change. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
LlmBindingConfig, LlmRouteConfig and LlmAuthorizationResult move to the exported config package so that a dialect implemented outside this module can resolve routes and authorize through the guard without reimplementing them. LlmBindingConfig no longer carries legacy dialect state, and authorize takes the credentials header name instead of a legacy dialect. LlmLegacyBindingConfig extends it with the legacy dialect resolver and signer for the legacy handlers. LlmDialectContext.attach takes the shared LlmBindingConfig, which the lean server and client factories build once per binding. The configuration module is required transitively because the config exposes its options type. No behavior change. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
Unit tests for a streaming decoder that turns an OpenAI Chat Completions request into typed blocks, written before the implementation against a skeleton that decodes nothing. Each case is decoded whole and fed in 1, 7 and 128 byte chunks. Covered: every section type, mixed content and a conversation, member order independence for role, model and tool_call_id, escaped and multibyte text, 10k and 100k blocks, a large verbatim object part, limited and exhausted sink availability, and rejection of a missing model or role, malformed or truncated or trailing input, and lookahead beyond the hold bound. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
A streaming decoder that turns an OpenAI Chat Completions request into typed blocks, built as function-pointer decoder states over a JsonParserEx in the style of the mcp server decoder. Text sections stream their decoded text as fragments. Tool calls, tool definitions and non-text content parts pass through as verbatim JSON objects captured from the input window. A tool result carries its tool_call_id as the block extension. Member order is not assumed. The model, and each message role and content part type, is found by scanning ahead over the bytes held for the request, bounded by the decode slot, and the request is rejected when the bound is reached first. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
Registers openai through the dialect SPI. The server handler decodes the request into typed blocks, opens the application stream once the model is known, and sends each block as its own init frame carrying the section type, the message index and the extension, with the payload fragmented across frames as the application window allows. A request that cannot be decoded is rejected with a reset, after acknowledging the bytes it discards. The reply path is unchanged and still forwards the native response. Known: the openai server ITs that read the request as native JSON on the application stream, and the openai round trip, no longer pass. They are migrated with the client side in following commits. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
Adds network and application script pairs for the openai request mapping beyond the per-section cases: tiny-frame fragmentation, a large verbatim object part, escaped and multibyte text, whitespace, ignored members, tools before model and messages, an unknown content part, and server-side rejection of a missing model, a non-object request, a missing role, non-array messages, malformed JSON and trailing content. Spec peer-to-peer ITs pass. The runtime ITs for these scenarios are added but not yet run, and the three encode ITs are expected to fail until the client side is implemented. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
Model detection happens before any section is decoded, so a request without a model is not a section rejection, and openai.invalid already covers it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
A request that is not an object is rejected at model detection, before any section is decoded, so it is not a section scenario. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
…r factory Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
…r payload Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
…ir payload Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
…t is rejected Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
…quest error Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
…is rejected Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
…cted Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
Replace the local FIN/INIT constants and bitmask tests with Flags so the openai server reads data frame flags the same way as the rest of the engine. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
…erver Test flags with Flags.hasInit and Flags.hasFin where they are used instead of capturing the results in local booleans, and build outgoing flags with a single sliceFlags helper. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
… the openai server Drop the sliceFlags helper. The decoded block keeps its pending INIT in a flags field instead of a boolean, and the other two sites derive the frame flags from their position with Flags.fin. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PiJJyeLu5ENYBaYCCJKJD2
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JpEmbPfbUtpL244vMsJHH5
…window An llm client now grants its request window only after the http client stream receives a window, so a server that cannot be reached resets the caller before any credit, which lets an llm proxy fall back to the next model. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JpEmbPfbUtpL244vMsJHH5
A fallback model is resolved against the proxy routes with the original dialect and authorization, retries on the exit it resolves to, and is skipped when no authorized route matches. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JpEmbPfbUtpL244vMsJHH5
…lback model Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JpEmbPfbUtpL244vMsJHH5
…y fallback Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JpEmbPfbUtpL244vMsJHH5
27768bf to
ae30376
Compare
…del override script a peer and a spec IT The proxy fallback scenarios are split into direct client and server pairs that run without the engine: the caller leg, the refused and accepted exit attempts, and the exhausted leg. The runtime proxy ITs compose them across two exits. The unreachable scenarios gain the missing application server and network client, and the server model override scenarios gain their application clients. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JpEmbPfbUtpL244vMsJHH5
A video content block has no place in openai or anthropic, so a client rejects it naming the section and the dialect pair, as for user-audio. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DuEfqFwmhwGviQ5CnhwFAr
…he server side The neutral application script now opens the window after the challenge, as the request scenarios do, and the network clients expect the encoded response without the dialect members. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DuEfqFwmhwGviQ5CnhwFAr
The request and response options transform scenarios show dialect frames dropped across a hop, so they are named for what they show. The response network pair is no longer symmetric, so it is exercised by the runtime ITs. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DuEfqFwmhwGviQ5CnhwFAr
Each section scenario now ends with the response of the same dialect, so the client and server ITs both carry the request and the response, instead of a rate limit reset that the rejected scenarios already cover. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DuEfqFwmhwGviQ5CnhwFAr
…n scenarios The decode scenarios become section scenarios named for what they carry (assistant text, tool call, reasoning, refusal, members, logprobs, usage, streaming), run from both the client and the server side. The streaming requests now ask for a stream, so the server replays the events. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DuEfqFwmhwGviQ5CnhwFAr
…rm response scenarios The cross-dialect response scenarios are dialect neutral, taking the request's dialect and the responder's dialect as properties, with one network script per dialect. They run from the server side, and from the client side where both dialects can express the response. The same-dialect encode scenarios are renamed into the section scenarios, and the duplicates are removed. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DuEfqFwmhwGviQ5CnhwFAr
…enarios The payload-native transform scenarios named the model after each dialect, so the other dialect's network script carried it. They now take the neutral model name, as the text transform scenarios do. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DuEfqFwmhwGviQ5CnhwFAr
…esponse The challenged model scenarios pair a network client that now expects the response with an application server that still reset the request. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DuEfqFwmhwGviQ5CnhwFAr
A section stream describes the same feature identically whichever dialect produced it, so the target wire for a feature is one wire rather than one per source dialect. The payload-native transform scenarios drifted apart: - inline images carry one media type and one data value - inline documents carry the same attributes, dropping the openai-only filename that the anthropic section never produces - document references use one file id and one scenario name - tool definitions carry the same description and are written before the user text - assistant tool calls share one scenario name, the same user text, message index and call ids, with two calls in the basic scenario - mixed content carries the same text and image sections Dialect-native error tails and rejection scenarios are unchanged. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DuEfqFwmhwGviQ5CnhwFAr
The transform scenarios shared by both dialects expected a different 429 reset depending on which dialect produced the section stream, which tied the target wire to its source. Every shared scenario now resets with status 429, type rate_limit_error and message "Rate limit exceeded", and the anthropic shaped network wires return the matching rate_limit_error body. Dialect-specific rejection and one-sided scenarios are unchanged. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DuEfqFwmhwGviQ5CnhwFAr
…response The model override server scripts answer with the response, so their client peers read it instead of a rate limit reset. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DuEfqFwmhwGviQ5CnhwFAr
…contexts
A metric context that evaluates attribute expressions can only read the
stream frames it observes. Expressions naming other configuration, such as a
guard referenced as ${guarded['name'].identity}, cannot be resolved because the
context receives no way to turn a configured name into a namespaced
identifier, which bindings obtain from their own config for this purpose.
MetricContext gains a supply overload that also receives the binding's name
resolver. It defaults to the existing attributes overload, so current metric
contexts are unaffected.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JmtXgg57sb1dimx5No9hHF
…ncomplete usage with attributes
Adds llm.interactions (counter), llm.time.to.first.token (histogram) and
llm.usage.incomplete (counter) to the llm metric group, and lets every llm
metric resolve telemetry attributes.
Attribute expressions: ${llm.status}, ${guarded['name'].identity} and
${guarded['name'].attributes.claim}. The status is the error status
reported on an aborted or reset exchange, 200 for a completed one, and is
omitted for metrics recorded before an exchange ends. Guarded expressions
resolve through the guard named in the expression, using the authorization
carried on the request begin.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JmtXgg57sb1dimx5No9hHF
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JmtXgg57sb1dimx5No9hHF
Attributes are limited to what a binding's telemetry.attributes can already express for a guard, so the interaction counter and the other llm metrics are labelled only by configured guarded expressions. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JmtXgg57sb1dimx5No9hHF
…zation The test binding relayed every frame it forwarded with the stream's authorization except FLUSH, which went out with authorization 0. A stream accepted for a non-zero authorization was then aborted when an advisory flush crossed the binding. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JmtXgg57sb1dimx5No9hHF
…ation The attribute test guard now configures credentials, so it resolves an identity and attributes only for a non-zero authorization. The openai.usage scenario takes its authorization as a property, defaulting to 0L, and the attribute IT overrides it to 1L, which fails if the request authorization is not passed to the guard. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JmtXgg57sb1dimx5No9hHF
ae30376 to
adf9005
Compare
|
CI on The failure is not from this PR. The Generated by Claude Code |
|
The re-run of The fix is in aklivity/zilla#2711 ( Generated by Claude Code |
…esidio analyzer (#2711) With gunicorn 25's control socket enabled, the analyzer's single worker intermittently never boots: the log stops after "Control socket listening" with no "Booting worker" line, the health probe times out, and compose reports the container unhealthy. Reproduced with the example's own compose file, 12 analyzers started in parallel: develop left 2 without a worker and unhealthy, with the control socket disabled all 12 were healthy. Claude-Session: https://claude.ai/code/session_01JmtXgg57sb1dimx5No9hHF Co-authored-by: Claude <noreply@anthropic.com> (cherry picked from commit 1b43d3d)
1be8ba8 to
440df46
Compare
Description
Adds three metrics to the
llmmetric group and lets everyllm.*metric label its series fromtelemetry.attributes.Stacked on aklivity/zilla#2681, which is stacked on #2680 and #2679. This branch is rebased onto #2681's head (
claude/gifted-cray-r71cze, currently98044e968), so until those merge the "Files changed" tab also shows their commits. Only the 6 commits after98044e968belong to this PR; once they merge, the branch rebases ontodevelopand the diff reduces to those.Metrics
llm.interactions(counter): one per completed, aborted or rejected exchange.llm.time.to.first.token(histogram, ns): request BEGIN to the first reply DATA frame.llm.usage.incomplete(counter): an exchange whose reply ends with no usage reported, or aborts after the reply opened.Attributes
telemetry.attributesnow apply to allllm.*metrics. The expressions are the existing guarded forms,${guarded['<guard>'].identity}and${guarded['<guard>'].attributes.<name>}, resolved against the named guard with theauthorizationcarried on the request BEGIN. The guard is looked up through the binding's own name resolution, so a claim a guard exposes as an attribute can label usage per tenant.llm-specific expression is added. The binding is already identified by the series' binding label.Engine
MetricContext.supply(recorder, attributes, resolveId)is a defaulted overload that also receives the binding's name resolver, so a context that evaluates attribute expressions naming other configuration (a guard) can resolve the name. It defaults to the existing attributes overload, so existing metric contexts are unaffected.FLUSHframes with the stream's authorization like every other frame it forwards. It sent them with authorization0, so a stream accepted for a non-zero authorization was aborted when an advisory flush crossed it.The
llm.proxyexample records the new metrics, the golden schema lists them, and the README documents the attribute expressions, including a JWT claim as a metric attribute.How this differs from #2587 as filed
The issue was written before the usage shape and the first metrics landed, so the delivered scope differs in these ways:
LlmFlushExcase. Usage is carried onLlmEndEx(a completed reply) andLlmAbortEx(an aborted one), so the partial-usage counter is driven by those frames.llm.tokens.*), not counters. A token count a dialect does not report is left unrecorded, never recorded as0, so a dialect reporting no usage still produces an interaction count with absent token fields.llmbinding that served the request.examples/llm.proxy/etc/test/verify.pydoes not assert the new metrics; I could not run the compose stack in this environment.Fixes #2587
Testing
MetricContextTest(the overload delegates),LlmMetricGroupTest(names, kinds, units, descriptions).LlmMetricsITruns thebinding-llm.specapplication scripts through atype: testproxy (thellmbinding is not loaded): usage (OpenAI and Anthropic, streaming and not), abort and rate-limit rejection, plus a guarded-attribute scenario. That scenario uses a test guard configured withcredentials, so it resolves an identity only for a non-zero authorization; theopenai.usagescripts takeproperty authorization 0Land the IT overrides it to1L. It fails if the request's authorization is not passed to the guard, and the exporter asserts thetenantanduserlabels on every series.MetricContextTestandGuardFactoryTest,metrics-llm.spec(7) andmetrics-llm(16, including the ITs) pass, with checkstyle, license headers and the JaCoCo gate on formetrics-llm. I have not re-runbinding-llm; the example jobs run in CI.🤖 Generated with Claude Code
https://claude.ai/code/session_01JmtXgg57sb1dimx5No9hHF