Repository navigation
[API Proposal]: Add a provider-neutral abstraction for decision-oriented AI models #7764
Description
Activity
Follow-up proposal in Microsoft Agent Framework: #8545 — Integrate Microsoft.Extensions.AI decision-model inference into agent decision points. That feature request is contingent on and builds on the provider-neutral abstraction proposed here.
@luisquintanilla There is now a working implementation of this proposed shape, name for name, with a provider and two consumers: microsoft/agent-framework#8563. It lives in Microsoft.Agents.AI.Abstractions as an experimental copy, intended to be deleted once MEAI ships the abstraction, plus a TypeSafe (Jev) provider package and an evaluation-framework consumer. Two small additions came out of building it that may be worth considering for the contract: a classified DecisionClientException with IsTransient, and construction-time range validation on the answer types. Happy to adjust to whatever shape API review lands on.
As a clarification, my proposal is to bring this implementation here as it is the natural place to be :) - say the word and a PR appears.
One detail worth making normative:
DecisionChoice.Nameshould be an opaque, caller-owned identifier, and providers should return the exact string value inSelectedChoiceand the probability keys without case or Unicode normalization. The human-readable criterion can remain inDescription.Dynamic action spaces make this important. In Jev Social, we rebuild candidates from each browser observation and hash the normalized operation plus arguments into stable IDs, then reject any selected value outside the exact criteria map before execution. This prevents an adapter rewrite or stale label-like result from becoming a different action.
If that is the intended contract, naming this property
Idor explicitly documenting exact round-trip identity would make independent adapters safer and more interoperable. I maintain Jev Social; sharing implementation evidence rather than proposing a provider-specific constraint.- added a commit that references this issue
on Sep 20, 2026 - addedarea-aiMicrosoft.Extensions.AI librariesMicrosoft.Extensions.AI libraries
on Sep 23, 2026 - added a commit that references this issue
on Sep 24, 2026 Hi folks,
Thanks for the interest and initial work here.
Given the OpenAI Decision API announcement, Ollama support, and other open-source models on HuggingFace, it makes sense for MEAI to expose an experimental API to support these capabilities
I have a few prototype implementations here for various providers - https://github.com/luisquintanilla/typesafe-meai/tree/1db913d1e26b4b5471e41897ea7b6c34915e2c15
One scenario which I don't think this original proposal covers is using an IDecisionClient as an AITool. That pattern is explored by MAF here - microsoft/agent-framework#8859
We'll post a PR with a formalized API surface and and would appreciate feedback.
cc: @rogerbarreto
Thanks, Luis. Here is some feedback from using decision models as evaluators in AgentEval. We have an
IDecisionClientof our own today; it is marked experimental so that it can adapt to whatever MEAI ships rather
than become a third abstraction.Five things mattered in practice:
- An undecided outcome distinct from "no". A timeout, a refusal to answer or an unparseable response should not
come back asfalse. In evaluation an undecided case has to leave the denominator, while a "no" counts against
the agent. A typed failure (rate-limited, invalid response, content-filtered…) also lets callers retry the right
ones. - The raw probability, not only a yes/no. The threshold is the caller's decision. On 298 labelled cases,
moving one threshold from 0.50 to 0.90 took false passes from 27.3% to 1.9% and false fails from 11.0% to
29.5%. A boolean-only API hides that trade-off. - A slot for reference context, separate from the state being judged. A short description of what is being
judged moved accuracy on one set from 64% to 96% with no false passes. Without a slot, callers paste it into
the state, which mixes "what to judge" with "what is being judged" and makes results harder to cache and
fingerprint. - Several questions about one state in a single call, with keyed answers. This should cover binary, choice and
score questions. An evaluator usually asks several criteria about the same response, and one call per question
multiplies cost and latency. - Provenance on the response. It should carry the resolved model and version (an alias such as
jev-latest
resolved to a concrete build) plus usage. Without it a result cannot be reproduced or compared across runs.
Happy to test the formal API PR against our labelled sets when it lands. Evidence (AgentEval repo):
- An undecided outcome distinct from "no". A timeout, a refusal to answer or an unparseable response should not
Thank you for the detailed proposal, @mo3in. We are going to look into this.
Reacted by Mo3in@luisquintanilla @jeffhandley — after reviewing the implementation work in [#7795 — Add experimental provider-neutral decision abstractions (Layer 1)] and [#7796 — Layer 2 decision interoperability], I think the current direction is strongly aligned with the intent of this original proposal.
The separation now looks particularly clean:
- [#7795] defines the provider-neutral decision capability itself:
IDecisionClient, Binary / Choice / Score observations, complete probability distributions, typedDecisionDefinition<TResult>binding, provenance, usage, protocol validation, source-generated JSON support, provider metadata, and classified provider failures. - [#7796] keeps composition deliberately small by integrating decisions with existing MEAI primitives rather than creating a parallel framework:
DecisionDefinition<TResult>.AsAIFunction<TState>()- stock
AIFunctionFactory - existing
RoutingChatClient - consumer-owned ingestion/routing/application policy.
That is, in my view, preferable to introducing separate decision-specific tool, router, task, or orchestration abstractions.
A few follow-up semantics seem worth recording here before the experimental contract settles. These are mostly based on scenarios that were open questions in the original proposal but now have implementation or consumer evidence behind them.
1. Reference context should be distinguishable from the state being evaluated
@joslat raised an important evaluator use case in [this comment]: a decision often depends on both the object being judged and separate reference material.
Conceptually, many calls are:
reference / policy / rubric / expected behavior + state being evaluated + questionsrather than simply:
state + questionsFor example:
Reference: "Refunds are allowed only for duplicate charges or cancelled orders." State: "The assistant approved a refund because the customer disliked the color." Question: "Is this response compliant with the refund policy?"Today this can mechanically be represented by embedding both values into
State, but they are semantically different inputs.Keeping them distinct can matter for:
- caching and fingerprinting,
- evaluator reproducibility,
- reusing one reference against multiple states,
- avoiding accidental mixing of instructions/reference material with the object under evaluation,
- provenance and dataset construction.
The AgentEval evidence shared by @joslat is especially interesting because providing reference context separately materially changed evaluation quality in their experiments.
I do not think we need to prescribe an API shape immediately, but it may be worth explicitly asking whether the portable request model should eventually support something conceptually like:
DecisionRequest { State Reference? // or Context? Questions }
The important semantic point is the separation, not necessarily the property name or representation.
2. Semantic abstention / undecided should be considered separately from failure
The current work has significantly improved failure semantics.
[#7795] now includes
DecisionClientExceptionwithIsTransient, and also distinguishes provider failures from invalid protocol responses.There is still another outcome worth discussing:
Answered Abstained / UnableToDecide FailedA model legitimately saying:
there is insufficient information to make this decision
is not equivalent to:
Binary = falseand is also not necessarily equivalent to:
provider failure timeout rate limit invalid protocol responseThis matters particularly for evaluation and safety-oriented consumers. For example, an evaluator should generally not turn "unable to determine whether the agent satisfied the criterion" into a negative judgment automatically.
It becomes even more important because
IDecisionClientsupports heterogeneous batches:Q1 → answered Q2 → answered Q3 → unable to decide Q4 → answeredThe original proposal left partial-result semantics open. With more concrete consumers now available, I think API review should explicitly distinguish:
request/provider failure ≠ protocol-invalid response ≠ valid semantic abstentionI would not necessarily add a
DecisionAnswerStatusor partial-result API yet. But I think the semantic requirement should remain visible so that the first experimental shape does not accidentally make abstention impossible to represent later.3. Capability discovery:
SupportedKindssolves the first layer, but probably not the whole problemOne of the original open questions was whether primitive support and provider limits should be discoverable.
[#7795] now makes useful progress here with:
DecisionClientMetadata.SupportedKinds
That is a good addition, especially now that provider proofs do not all expose identical capabilities.
For example, the current implementation evidence already includes providers with different supported decision kinds.
The remaining question is whether any additional constraints eventually become sufficiently portable to expose through common metadata, such as:
maximum questions per request maximum Choice candidates maximum Score levels batch support / limits native scalar Score support input/state size limits provider probability precisionI would not put all of these into v1.
Rather, I suggest refining the original open question from:
Should supported primitive kinds and provider cardinality limits be discoverable?
to something closer to:
Which execution capabilities and limits, beyond
SupportedKinds, are sufficiently portable and useful to expose through common decision-client metadata?That lets
SupportedKindsbe considered a solved part of the proposal while keeping the abstraction extensible for real provider constraints.4. The exact opaque identity semantics introduced in Layer 1 look important and should remain normative
@IRONICBo previously pointed out the importance of candidate identity for dynamic action spaces in [this comment].
The current Layer 1 shape has moved in a good direction here:
DecisionCandidate.Id DecisionProbability.Id ChoiceDecisionAnswer.SelectedCandidateId
rather than treating candidate text as identity.
I think this exact round-trip behavior should remain part of the semantic contract:
caller-owned candidate ID → provider evaluation → exact same candidate ID returnedwith ordinal/string identity semantics rather than provider-side case folding, Unicode normalization, label rewriting, or semantic matching.
This is important not only for classification but also for scenarios such as:
tool shortlisting dynamic actions agent handoffs workflow branches model routing skill selectionwhere a selected value may ultimately map to executable application behavior.
Relationship to Layer 2 / AITool
The AITool scenario @luisquintanilla called out earlier is now addressed cleanly by [#7796].
In particular, I like that the implementation does not introduce a new decision-tool framework.
Instead:
DecisionDefinition<TResult> .AsAIFunction<TState>(...)
is only a thin convenience over the existing
AIFunctionFactory, while applications needing custom behavior can simply expose an ordinary annotated method.Likewise, decision-based model routing remains a consumer of:
RoutingChatClient.Create(...)
rather than becoming a new decision-specific router.
That maintains the architectural distinction from the original proposal:
IDecisionClient = provider/model inference capability AIFunction / RoutingChatClient / evaluators / agents = consumers of that capabilityThe related Agent Framework work also reinforces that separation; see the AITool exploration in [microsoft/agent-framework#8859].
Summary
With [#7795] and [#7796], most of the core goals from this proposal are now represented much more concretely:
provider-neutral probability-preserving heterogeneous-question aware typed when useful AOT/source-generation friendly strict about protocol correlation composable with existing MEAI primitives independent from agent orchestrationThe main semantic areas I would keep visible during API review are therefore:
- optional reference/context distinct from evaluated state;
- semantic abstention / undecided distinct from provider or protocol failure;
- future capability discovery beyond
SupportedKinds; - exact opaque identity round-tripping for candidates/levels.
None of these require expanding Layer 2 into another framework. They are primarily questions about making the underlying provider-neutral decision contract robust enough for evaluation, routing, dynamic action spaces, and other consumers that are already beginning to appear.
- [#7795] defines the provider-neutral decision capability itself:
One additional provider-neutral concern seems worth considering before the
DecisionRequestinput contract settles: multimodal decision input.The current Layer 1 proposal in [#7795] models the shared decision state as:
DecisionRequest( JsonElement state, IReadOnlyList<DecisionQuestion> questions)
This is a strong representation for structured application state, and the typed
TState + JsonTypeInfo<TState>helpers provide a useful AOT-friendly application boundary.However, decision-oriented inference is not necessarily limited to JSON-shaped state.
Recent provider work is already exploring decision evaluation over inputs such as text and images. A provider-neutral abstraction may therefore eventually need to evaluate:
structured application state text image document/data content or combinations of themFor example:
Structured state: claim type = shipping damage customer tier = gold Supporting content: photograph of the received package Questions: Is visible damage present? Which damage category applies? How severe is the damage?Encoding image or document content inside
JsonElementwould effectively turnStateinto a generic transport envelope, which seems undesirable.MEAI already has reusable content abstractions such as
AIContent,TextContent, andDataContent, so I think API review should consider whether decision inference should eventually compose with those existing primitives rather than introduce a separate multimodal representation.I am not suggesting replacing structured
StatewithIReadOnlyList<AIContent>.Structured state and supporting content are semantically different. Conceptually, the model may instead be closer to:
Decision input ├── structured State │ └── JsonElement / typed TState │ ├── optional supporting Content │ ├── text │ ├── image │ └── other MEAI content │ └── QuestionsThis also relates to the separate reference/context discussion, since reference material may itself eventually be multimodal.
The exact public API does not need to be decided now, but I think API review should explicitly answer:
Is
JsonElementintended to define the complete portable input domain ofIDecisionClient, or should the abstraction leave a first-class path for multimodal MEAI content alongside structured state?Because [#7795] defines the foundational provider boundary, this seems easier to account for now than after providers begin treating
Stateas necessarily JSON-only.This should not expand the scope of [#7796]; the
AIFunctionand routing composition layer can remain unchanged.
Background and motivation
Microsoft.Extensions.AIprovides provider-neutral abstractions for several distinct AI model capabilities, including chat completion, embeddings, image generation, speech-to-text, text-to-speech, and realtime interaction.A new category of AI models is emerging whose primary operation is different from text generation:
The motivating implementation is TypeSafe AI's Jev, described as a "System One Model".
Jev is not being proposed as the abstraction itself. It is an example of a model exposing a capability that currently does not have a corresponding provider-neutral contract in
Microsoft.Extensions.AI.The important distinction is between generative inference:
and decision-oriented inference:
The output space of each question is defined before inference. Applications consume probabilities, distributions, selections, and scores directly rather than parsing generated prose.
A concrete example is:
These values can participate directly in normal application code:
This proposal asks whether
Microsoft.Extensions.AI.Abstractionsshould define a small provider-neutral capability for this form of inference.The exact naming is open to API review. For the proposal below, I use
IDecisionClient.Why this is not just structured output from
IChatClientAn
IChatClientcan certainly be prompted to return JSON or use a structured response format.That provides syntactic structure, but does not define the semantics of a decision-model capability.
A decision-oriented contract can standardize concepts that
IChatClientintentionally does not:For example, an application should be able to ask for a categorical choice and rely on the response containing a probability distribution over exactly the supplied candidates.
With
IChatClient, that behavior is an application-specific prompt/schema convention.With a dedicated decision-model abstraction, it becomes part of the model capability contract and can be implemented by different providers.
This would be analogous to the reason embeddings have
IEmbeddingGenerator<TInput,TEmbedding>rather than being represented as "ask anIChatClientto output an array of floats".Proposed semantic model
The initial abstraction should support three general-purpose decision primitives.
The names below are intentionally provider-neutral.
BinaryDecisionQuestionChoiceDecisionQuestionScoreDecisionQuestionThese map naturally to Jev's current Noul, Choice, and Score capabilities, but none of those provider-specific names need to appear in the common API.
Binary decision
A binary decision answers a yes/no proposition.
The result is:
A value near
1means strong support for true, a value near0means strong support for false, and a value near0.5represents uncertainty between the two outcomes.There should not be a separate mandatory
Confidenceproperty for this primitive.The probability already captures the two-outcome distribution:
This also matches Jev's Noul semantics, where no separate confidence value is returned.
Optional true/false criteria should be supported so callers can clarify the semantic boundary:
Choice decision
A choice selects exactly one member of a caller-supplied unordered set.
For example:
The result should contain:
Probabilitiesrepresents the complete categorical distribution and should contain an entry for every supplied choice.The probabilities should each be in
[0, 1]and sum to approximately1.SelectedChoicemust be one of the supplied choices.Confidenceshould be optional in the common abstraction. Some providers, including Jev, expose a useful scalar derived from the shape of the distribution, but the full probability distribution is the more fundamental interoperable value.The contract should not claim that confidence is calibrated or comparable across providers unless an implementation explicitly documents that guarantee.
Score decision
A score represents a position on an ordered caller-defined set of levels.
For example:
Unlike a Choice, these values have an order.
The result should contain:
The probability distribution represents the model's probability for every supplied level.
For an
N-level scale, levels have ordinal positions:The continuous score can then have a well-defined provider-neutral meaning:
For example:
This is useful because two results with the same rounded level may still carry materially different distributions.
As with Choice,
Confidenceshould be optional and should not replace the full distribution.Batching heterogeneous questions is a core requirement
A central property of this capability is that multiple questions may be evaluated against the same state in one model request.
For example:
The request must therefore support a heterogeneous collection of question types.
This should not require callers to make one request per decision.
A provider may optimize these questions using parallel inference, shared state encoding, batching, or some other mechanism.
The abstraction should expose batching as a first-class capability without prescribing how the provider implements it internally.
An important semantic property is that every question is evaluated against the same supplied state.
Providers may have limits on:
Those limits should remain provider/model-specific rather than being encoded from Jev into the common abstraction.
API Proposal
The following is intended as a concrete starting point for API review, not a claim that every name is final.
Request
JsonElementis proposed for the shared state because the state may naturally be:and because it avoids requiring providers to serialize arbitrary runtime
objectinstances.It also provides a predictable boundary for trimming and Native AOT scenarios.
A convenience API could later allow strongly typed state:
This helper could serialize the state with the supplied
JsonTypeInfo<TState>without making the provider-facing abstraction generic.Question model
Question IDs must be unique within a request.
They are application correlation identifiers and should not be assumed to carry model semantics.
Binary question
Choice question
Choice names must be unique within the question.
Their descriptions provide semantic criteria but do not impose ordering.
Score question
A Score must contain at least two levels.
Their list order defines their ordinal position:
The abstraction should not adopt Jev's current maximum number of levels as a general API constraint.
Response model
The key in
Answerscorresponds to the question ID.The answer types are heterogeneous:
Binary answer
Invariant:
Choice answer
The common semantic expectation is:
Score answer
For levels indexed
0 ... N - 1:and the portable Score semantic is:
A provider adapter may compute
Scorefrom its probability distribution if the underlying provider does not return the expected value directly.Options and metadata
The capability should follow existing MEAI conventions where useful.
A minimal options type could be:
And metadata could be available through
GetService:This mirrors existing MEAI capability patterns without requiring a larger builder/middleware proposal yet.
Example usage
Mixed decision batch
Binary probability
Choice distribution
Continuous Score
The important point is that the provider returns decision values, while application code remains responsible for thresholds and business actions.
Why probability is part of the contract
I think probability should be treated as part of the semantics of these primitives rather than hidden in
AdditionalProperties.For example, this API:
would lose important information compared with:
Similarly:
loses information compared with:
That distribution can materially change how an application behaves.
The same applies to Score. A score of
1.0might mean:or:
Those have the same expected score but very different uncertainty.
For that reason, the full distribution should be the portable result for Choice and Score.
Confidence, in contrast, can remain optional because it is a summary statistic derived from the distribution and its exact computation may vary between providers.Probability and calibration semantics
The abstraction should distinguish probability from calibration guarantees.
A probability value means:
The common API should not imply:
Calibration is a property of a particular model/provider and possibly of a particular workload.
Providers that make calibration guarantees can expose that through metadata, documentation, or future capability metadata.
Likewise,
Confidenceshould not be assumed comparable between providers unless explicitly documented.State and AOT considerations
I do not think the provider-facing API should accept:
because that leaves serialization policy implicit and makes Native AOT/trimming harder.
JsonElementprovides a simple stable interchange representation.A strongly typed convenience overload accepting:
can preserve ergonomic and source-generated serialization without requiring the core client abstraction to become generic.
This also allows the same request to naturally carry either structured state or a simple JSON string.
Structured instructions and criteria
Some providers may support richer instructions or criteria than plain strings, such as JSON objects or arrays carrying examples, positive/negative guidance, or provider-specific hints.
I would not standardize those richer shapes in the first version unless there is evidence they are common across providers.
The portable v1 contract can use textual instructions/descriptions.
Provider-specific richer representations can remain reachable through:
This keeps the common abstraction meaningful rather than turning it into a generic JSON transport.
Failure and capability semantics
There are a few points where maintainer feedback would be especially useful.
The initial API could use normal request-level exceptions when an implementation cannot execute the request.
However, heterogeneous batches raise useful future questions:
I would prefer not to lock these into the initial proposal without evidence from multiple providers.
The core API should nevertheless be designed so such capability metadata can be added later without redesigning the basic request/answer model.
Relationship to
IEvaluatorMicrosoft.Extensions.AI.Evaluation.IEvaluatorrepresents an evaluation operation/framework.A decision client would represent an inference provider capability.
These can compose:
For example, an evaluator might use a Score or Binary decision internally.
That does not make the two abstractions equivalent.
The same distinction exists between an evaluation framework and the
IChatClientit may currently use.Relationship to reranking
A reranker answers a much narrower question:
A decision model covers general application judgments such as:
A reranker could potentially be implemented using a decision model, but a general decision client should not depend on retrieval-specific concepts such as documents, search queries, or ranking pipelines.
This is therefore orthogonal to the retrieval/reranking discussion in
dotnet/extensions#7507.Relationship to chat routing
This proposal is also separate from
RoutingChatClientand the response-quality/cascading discussion indotnet/extensions#7712.Those abstractions answer:
A decision client answers:
A router could choose to consume an
IDecisionClientas part of its policy, but routing is a consumer of decision inference rather than the decision-model abstraction itself.Why not one completely generic structured-inference API?
A possible alternative is something such as:
or:
That would be flexible, but it would standardize very little.
Two implementations could satisfy such an interface while having no interoperable semantics.
The proposed Binary / Choice / Score primitives deliberately standardize a small set of useful decision semantics:
That gives middleware and consuming libraries something meaningful to build on.
Why not one interface per primitive?
Another option would be:
That would make individual primitives strongly typed, but it makes heterogeneous batching around one shared state difficult and can force providers to make multiple calls even when their native API handles all question types together.
A single
IDecisionClientwith typed question and answer subclasses better represents providers that natively evaluate mixed decisions in one request.Why not a generic
IDecisionClient<TState, TResult>?This could improve compile-time typing for one decision shape, but it does not naturally represent this important request:
A generic result type therefore works against heterogeneous batching.
The non-generic provider contract plus typed question/result primitives appears to provide a better interoperability boundary.
Typed helpers can be layered above it.
Existing related proposals
Before filing this proposal I reviewed several adjacent API discussions.
dotnet/extensions#7507discusses retrieval pipeline abstractions andIReranker. That is retrieval-specific and narrower than general decision inference.dotnet/extensions#7712discussesCascadingChatClientand intentionally keeps response-quality policy outside the routing client. A decision model could become one possible implementation of such a policy, but it is not itself a routing client.dotnet/extensions#7587is useful precedent for deciding whether a capability belongs inMicrosoft.Extensions.AIversus a separate peer package. I think that same question is worth explicit maintainer feedback here.Package placement
My initial hypothesis is:
because this is a model-provider capability analogous to:
rather than an application workflow domain.
However, I would like maintainer feedback on this point.
The important objective is not the package name; it is establishing the right provider-neutral boundary if this category is considered sufficiently general.
Evidence and ecosystem maturity
I want to be explicit about an important limitation of this proposal.
Jev is currently the clearest concrete motivating implementation I have found for this exact combination of:
There are many adjacent technologies—classifiers, reward models, judge models, semantic routers, guardrail models, and rerankers—but they should not be claimed to implement the same capability without verifying their semantics.
So this issue is not claiming that a mature multi-provider ecosystem already exists.
Instead, the question is:
If the maintainers believe one provider is not enough evidence yet, an experimental API or further prototype/provider survey may be the appropriate next step.
Design goals
The proposed abstraction should be:
It should not attempt to standardize:
Those can be built above the core model capability.
Potential consumers
A provider-neutral decision capability could later be used by libraries and applications for scenarios such as:
These are examples of consumers, not responsibilities of
IDecisionClient.In particular, model output should not be treated as authorization:
not:
Follow-up integration
If MEAI adopts a provider-neutral decision-model abstraction, a separate proposal can explore integration with Microsoft Agent Framework.
Potential consumers there include:
That integration should depend on the MEAI abstraction rather than defining a competing Agent Framework-specific decision-model interface.
This issue is intentionally scoped only to the underlying model capability.
Open questions for API review
The areas where I would especially value maintainer guidance are:
IDecisionClientthe right conceptual name, or is another term more appropriate?Microsoft.Extensions.AI.Abstractions, or should it be a peer abstraction package?JsonElementthe appropriate provider-facing representation for shared state?Confidenceexist in the common contract at all, or should consumers always derive their own certainty measure from probabilities?TState + JsonTypeInfo<TState>helpers ship together with the abstraction?The main goal of this proposal is to determine whether decision-oriented inference deserves a first-class provider-neutral model capability in the Microsoft.Extensions.AI ecosystem, in the same way that applications today can depend on
IChatClientorIEmbeddingGeneratorwithout depending directly on a specific model provider.