You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Add agent-loop orchestration / toolMode: 'auto' to scope.models #612
Add agent-loop orchestration / toolMode: 'auto' to scope.models
Context
The model-access API (#510) declares toolMode: 'return' | 'auto' on GenerateOpts. 'return' mode is trivial (backend returns tool-call requests; caller resolves them externally). 'auto' mode is the in-process agent loop — scope.models resolves model-issued tool calls against scope.resources and registered MCP tools, executes them, and re-invokes the backend until the model produces a terminal answer.
That loop is substantial enough to deserve its own issue: tool resolution conventions, MCP integration surface, streaming behavior across tool-call boundaries, audit, error handling, and safety limits all need design work. Implementing it inside #510 would gate the model-access primitives behind 1-2 weeks of agent-loop design. This issue carries that work separately.
Once shipped, toolMode: 'auto' is what unlocks Harper's compressed-stack agent story: tool calls resolve in-process against local Resources, audited via the standard transaction log. The loop itself is the same regardless of target; the invocation latency depends on the tool target — sub-millisecond for Resource-backed tools (in-process dispatch), normal HTTP for tools the unified registry resolves from external MCP servers. The compressed-stack win applies to the Resource-backed case.
If the result has no tool calls, return to caller. Done.
If tool calls are present: validate each call's arguments against the tool's JSON Schema (opts.toolArgValidation: 'strict' | 'lenient' | 'none', default 'strict' — models hallucinate JSON shape often enough that this catches real bugs cheaply). Resolve each call against scope.resources (or the unified MCP tool registry), execute, capture results.
Append the assistant's tool-call message and the tool results to the message list.
Re-invoke the backend with the updated message list.
Repeat from step 3 until terminal (no more tool calls) or any safety limit trips.
Safety limits
The orchestrator enforces three independent limits, any one of which trips a structured budget-exceeded abort with the partial conversation for debugging:
opts.maxToolIterations — bounds the count of tool-call rounds. Default 10.
opts.maxTokens — bounds the cumulative prompt + completion + tool-output tokens across the whole loop. No default (uncapped). Lets dependent issues (e.g. Built-in Harper Agent Component harper-pro#676's built-in agent enforcing per-session token budgets) bound spend at the orchestrator rather than via post-hoc inspection of analytics.model_call.
opts.maxCostUsd — bounds cumulative USD cost. Per-call cost is computed from analytics.model_call.gpu_ms / token counts × the configured rate card (per-model). No default (uncapped). Resolves the per-session cost-cap ask from Built-in Harper Agent Component harper-pro#676.
A single iteration with a large prompt + large tool result + large response can blow a budget on its own — count is not a proxy for spend.
Parallel tool calls
Both OpenAI and Anthropic backends emit multiple tool calls in a single assistant message. Default execution: parallel — matches model expectations, latency-optimal. Configurable via:
opts.toolParallelism: 'parallel' | 'serial'
Within a parallel batch, mutations do not share a transaction boundary — each tool call is its own Resource dispatch and commits independently (see "Transaction scope per turn" below). Apps that need atomicity across multiple steps should wrap the multi-step intent in a single Resource method exposed as one tool, not rely on parallel-batch atomicity.
Tool result size cap
A tool returning 10MB inflates the next round-trip's prompt and can exceed the model's context window. The orchestrator enforces a per-tool-result size cap with a truncation marker:
opts.toolResultMaxBytes — default 64KB. Result exceeding the cap is truncated and a marker (e.g. […truncated; full result <Nbytes> bytes]) is appended; the truncated form is what feeds back into the model.
Cheap to add now; awkward to retrofit once apps depend on full results flowing through.
AbortSignal propagation across the loop
When the caller aborts a generate mid-loop:
Cancel the upstream LLM call in flight.
Propagate the abort into in-flight tool executions via the tool's invocation context (AbortSignal on the dispatched Resource call, when supported).
Already-committed writes from completed tool calls stay committed — no implicit rollback. The orchestrator returns a structured abort error.
Important for long-running agent endpoints (e.g. SSE-served chat completions where the client disconnects).
Tool resolution
Once the native MCP server in #465 lands, the scope.resources registry and "registered MCP tools" collapse into one unified tool registry — the same set of tools is exposed to external MCP clients (via #618 Application profile) and to in-process scope.models orchestration.
Sources of tools in 'auto' mode:
Auto-generated from the Resources registry via #618. Each @export-ed Resource produces get_*, search_*, create_*, update_*, delete_* tools (only the verbs the class implements, RBAC-filtered).
Custom Resource methods opted in via static mcpTools declaration per #622. For non-verb methods or computed tools.
Tool resolution is a lookup in the central tool registry from #615, which already handles RBAC-aware filtering. scope.models consults the same registry external MCP clients see.
Tool name convention
Inherited from #618: tool names are <verb>_<resource> (get_product, create_order, etc.) with sanitization for path separators. This orchestration uses the same naming so external MCP clients and in-process scope.models callers reason about the same set of tools.
Auto-discovery via the unified registry
With #618 landed, every @export-ed Resource is automatically a candidate tool — no explicit registration in GenerateOpts.tools required. The caller can still narrow the exposed set per call (tools: ['search_product', 'get_product']) or pass an explicit list for tighter control. Default behavior in 'auto' mode: the full RBAC-filtered registry the calling user can see.
Streaming with tool calls
For generateStream with toolMode: 'auto':
Yields content deltas as they arrive from the backend.
When the model emits a tool-call delta, the iterator yields that delta (signaling the consumer; useful for UI "agent is calling X" indicators).
The orchestrator suspends the upstream stream, executes the tool call(s), feeds results back to the backend.
The new stream from the backend resumes, yielding more deltas.
Loop until terminal.
The existing GenerateChunk shape (deltaContent?, deltaToolCalls?, finishReason?) supports this transition without changes.
Audit and replication
Tool calls inherit Harper's standard machinery for free:
Each tool call goes through the Resource dispatch — auth chain runs (tool calls have the caller's permissions), the call is logged in the transaction log, mutations replicate.
The model call itself is logged in analytics.model_call (per Add unified model-access API (scope.models) #510). The model call records the tool-call count; individual tool calls show up in the transaction audit log alongside data changes.
Querying "what did this agent actually do" is a regular Harper query against the audit log filtered by request / conversation / tenant.
This is the compressed-stack advantage made concrete: an agent's actions are recorded the same way as a human user's actions, in the same store, queryable with the same primitives.
Transaction scope per turn
Each tool call commits its own transaction. If the agent calls create_order and then update_inventory and the second fails, the first is already committed. This is the right default for composability with non-agent callers (tools-as-Resources behave identically whether called by an agent or directly), but worth calling out explicitly:
Multi-step atomic intent should be wrapped in a single Resource method (exposed as one tool), not relied on across multiple tool calls.
The orchestrator does not provide automatic rollback or saga semantics across the loop.
ConversationResource integration
Optional affordance for callers using ConversationResource (#511): pass opts.conversation?: ConversationResource and the orchestrator appends each turn (user message, assistant message, tool calls, tool results) as it runs. Without it, conversation persistence is the caller's responsibility — they'd reimplement the same write pattern around every agent loop invocation.
The orchestrator does not require a ConversationResource. Stateless agent calls work fine. When provided, the integration:
Appends the input as a user/system turn before the loop starts.
For each tool call: appends the assistant's tool-call turn and the tool result(s) as a role: 'tool' turn.
Appends the final assistant response as a turn before returning.
The contract: the orchestrator writes turns as it runs; the caller decides whether to provide a conversation at all.
Error handling per tool call
When a tool call fails, default behavior: append the error as the tool result (so the model can react and choose to retry, work around, or abort). Configurable via opts.toolErrorMode:
'recover' (default) — return error to model, continue loop
'abort' — abort the loop, return error to caller
Return shape
On success, scope.models.generate({ toolMode: 'auto' }) returns the final assistant response — same GenerateResult shape callers get from 'return' mode, just without toolCalls (those were resolved internally). Low overhead; matches the "I just want the answer" common case.
Opt in to the full sequence (assistant turns, tool calls, tool results, intermediate responses) via opts.includeToolTrace: true. Returns the response plus an ordered trace: ToolTraceEntry[] for inspection and debugging.
On error or safety-abort (budget exceeded, abort signal, tool error with 'abort' mode), the trace is included regardless of includeToolTrace. Operators need that trace to diagnose; hiding debug info on failure is an antipattern.
Acceptance
Audited against origin/main 2026-08-06. This list had gone stale — every box was unchecked despite #848 (merged 2026-06-02) implementing most of the loop. Verdicts below are from reading resources/models/agentLoop.ts, resources/models/types.ts, and HarperFast/documentationreference/models/tool-calling.md, not from commit messages. 9 done, 2 partial, 5 not started — the two partials (budgets, docs) are left unchecked.
The five not-started items are not five independent pieces of work: four of them (resource resolution, MCP resolution, transaction-log audit, permissions) all block on the same missing seam — dispatching to a Resource instead of a caller-supplied handler — which is #1740. The fifth is a JSON Schema validator decision.
opts.maxToolIterations, opts.maxTokens, and opts.maxCostUsd independently enforced; any trip returns a structured budget-exceeded abort with the partial trace. (Partial.maxToolIterations is a hard bound and enforced. The token/cost caps shipped as maxToolTokens / maxCostUsd — note the rename from maxTokens above — and are best-effort on the sync path: they derive from backend usage, so a backend reporting none warns once and continues rather than silently no-opping. Both throw 501 on generateStream because GenerateChunk carries no usage in v1. maxCostUsd additionally needs a rate card — the v1 cost function returns 0, so it only trips if a non-zero function is wired. BudgetExceededError.partialTrace is in place.)
Parallel tool-call execution by default; opts.toolParallelism: 'serial' runs serially. (Default 'parallel' via Promise.all; each handler is its own dispatch, no shared transaction.)
Per-tool-result size cap (opts.toolResultMaxBytes, default 64KB) with truncation marker. (Default 65_536; the model sees the truncated form, the trace records the original size.)
AbortSignal propagation: caller-side abort cancels upstream LLM call AND in-flight tool executions; already-committed writes stay committed. (Loop-level AbortController composed with the caller's signal; the composed signal reaches both the inner generate and ToolHandlerContext.signal. Budget trips fire loopController.abort() so an in-flight LLM call cancels rather than being awaited.)
Transaction scope per turn explicitly documented — each tool call commits independently. (Stated on toolParallelism in types.ts and in the tool-calling doc.)
Streaming + tool calls: chunks yield correctly through tool-execution transitions. (finishReason is treated as an internal signal — stripped from forwarded chunks and re-emitted as one terminal chunk after the assistant turn is persisted, so a consumer stopping on finish-reason cannot race past the conversation append.)
Error handling per opts.toolErrorMode; trace returned on error/abort even when includeToolTrace would otherwise be off. ('recover' default vs 'abort'; 'abort' throws before the conversation sink sees the round's tool turns, so the store never holds an error turn the model never consumed.)
Tool argument validation strict by default; configurable per opts.toolArgValidation. (Not started, and the shipped default inverts this line: v1 implements 'none' as the default; 'strict' and 'lenient' are reserved on the type surface and throw 501 at loop entry. Needs a JSON Schema validator decision — Harper uses Joi internally and passes JSON Schema through to backends, so adopting Ajv is its own call. Either build it or amend this criterion to match the shipped default.)
Per-tenant accounting on the model calls is preserved through the loop (each round-trip writes its own analytics.model_call row). (The loop calls back through models.generate(..., {toolMode: 'return'}) per iteration so each backend round flows the single-shot path and writes its own hdb_model_calls row; the outer 'auto' call stays out of the table.)
Dependencies
Hard: model-access API (Add unified model-access API (scope.models) #510), specifically Phase 3 — needs a backend that natively supports tools (openai backend, then anthropic / bedrock follow). ollama does not currently expose tools in a portable way; if/when it does, this orchestration picks it up automatically.
Hard: native MCP server family (#465 umbrella), specifically:
#615 — Tool registry + class-level introspection + RBAC-aware filtering (the registry this orchestration consults)
Previously this dependency was framed against the external HarperFast/mcp-server addon; that path is superseded by the native MCP server (and HarperFast/mcp-server is being archived per #623).
Pairs with Add ConversationResource for agent memory and conversation state #511 (ConversationResource) — turns.tool_calls and turns.tool_results record what the orchestrator did; optional integration via opts.conversation per the "ConversationResource integration" section above.
Downstream consumer: Built-in Harper Agent Component harper-pro#676 (Built-in Harper Agent Component) — uses this orchestration as its loop; surfaced the per-call cost-cap and per-turn cost-surfacing asks that informed the safety-limits section.
Out of scope
Tool catalog / discovery UX in Studio.
Cross-tenant tool invocation.
A full multi-agent orchestration layer (this issue does single-agent tool-call loops; multi-agent coordination is its own design space).
Per-model-call permission scoping (opts.toolPermissions to run the loop with a tighter permission set than the caller's). Tool calls inherit the caller's permissions in v1; tighter scoping is a v1.1 / follow-up consideration.
Structural cycle detection.maxToolIterations already catches infinite loops eventually. A "same tool, same args, N times consecutively" check would fail faster with a more actionable error — fast-follow if needed; not load-bearing for v1.
This looks great! Let's move forward with this!
Notes from Claude conversation:
Foundation looks solid — the loop semantics, RBAC-via-Resource-dispatch, and unified-registry anchor are all the right calls. A few minor gaps / adjustments worth folding in (none are architectural; mostly defaults and safety limits that are cheap to add now and awkward to retrofit):
Cost cap, not just iteration cap.maxToolIterations bounds count, not spend. One iteration with a large prompt + large tool result + large response can blow a budget on its own. Add opts.maxTokens and/or opts.maxCostUsd per generate call, with a structured budget-exceeded abort error. This also lets dependent issues (e.g. the built-in agent in Built-in Harper Agent Component harper-pro#676) enforce session budgets through the orchestrator rather than externally polling analytics.model_call.
Parallel tool calls — explicit semantics. Both OpenAI and Anthropic backends emit multiple tool calls in a single assistant message. The proposal doesn't specify execution order. Recommend parallel by default (matches model expectation, latency-optimal) with opts.toolParallelism: 'parallel' | 'serial' opt-out, and document that mutations within a parallel batch don't share a transaction boundary spanning the batch.
Transaction scope per turn — explicit decision. Each tool call goes through Resource dispatch and commits its own transaction. If the agent does create_order then update_inventory and the second fails, the first is already committed. That's almost certainly the right default (composability with non-agent callers), but worth calling out explicitly in the acceptance criteria so app authors know to wrap multi-step atomic intent in a single Resource method (called as one tool) when they need atomicity.
AbortSignal propagation. Spec what happens when the caller aborts a generate mid-loop: (a) cancel the upstream LLM call, (b) propagate abort into in-flight tool executions via context.abortSignal, (c) already-committed writes from completed tool calls stay committed (no implicit rollback). Important for long-running agent endpoints.
Tool result size cap. A tool returning 10MB inflates the next round-trip's prompt and can exceed the model's context window. Recommend a default per-tool-result size cap (e.g. ~64KB) with truncation marker, configurable via opts.toolResultMaxBytes. Cheap to add now; painful to retrofit once apps depend on full responses flowing through.
Tool argument validation default (open decision Add a CODEOWNERS file to set PR reviewers #5). Recommend strict by default — models hallucinate JSON shape often enough that this catches real bugs cheaply. opts.toolArgValidation: 'strict' | 'lenient' | 'none' for opt-out.
includeToolTrace default (open decision Default to attempting to serve index.html for a path #4). Recommend: final response only on success (as proposed), but always include the trace on error / safety-abort. Defaults that hide debug info on failure are an antipattern — operators need that trace to diagnose.
ConversationResource integration boundary. "Pairs with Add ConversationResource for agent memory and conversation state #511" is mentioned, but the contract isn't specified: does the orchestrator write turns to a ConversationResource as it iterates, or is conversation persistence the caller's responsibility? If it's caller-side, every app reimplements the same persistence loop. Recommend a small affordance: opts.conversation?: ConversationResource that the orchestrator appends to as it runs. Out of scope to require it, but worth supporting.
Latency claim precision. The "sub-millisecond invocation latency vs. MCP-over-HTTP" framing holds for Resource-backed tools (in-process dispatch) but not for tools the unified registry resolves from external MCP servers. Worth being precise — the loop is the same; the tool target determines latency. Otherwise a benchmark against an MCP-backed tool will appear to contradict the claim.
(Minor) Cycle detection.maxToolIterations catches loops eventually; a structural cycle check ("same tool, same args, N times in a row") fails faster with a more actionable error. Optional, low priority.
Most of these can land as additions to the existing "Open decisions" / "Acceptance" sections rather than new design work. Happy to PR a body edit if useful.
Parallel tool calls subsection: parallel by default, opts.toolParallelism: 'serial' opt-out, with the explicit note that a parallel batch doesn't share a transaction boundary.
Tool result size cap subsection: opts.toolResultMaxBytes default 64KB with truncation marker.
AbortSignal propagation across the loop: explicit semantics for upstream cancel + in-flight tool cancel + committed-writes-stay-committed.
Transaction scope per turn explicit note under Audit: each tool call commits independently; atomic intent belongs in a single Resource method.
ConversationResource integration section: optional opts.conversation — orchestrator appends turns when provided; stateless calls still work.
Latency claim tightened — sub-millisecond applies to Resource-backed tools (in-process dispatch), not to tools the registry resolves from external MCP servers. The loop is the same; tool target drives latency.
toolArgValidation folded into step 4 of the loop semantics: strict by default; 'strict' | 'lenient' | 'none' opt-out.
includeToolTrace folded into a new "Return shape" section: final on success; full trace on opt-in; trace always included on error or safety-abort.
Cycle detection + per-model-call permission scoping moved to Out of scope as fast-follow / v1.1.
Acceptance checklist updated to match. Open-decisions section dropped — decisions live in the body now, future work in Out of scope.
@kriszyp I broke this out into sub-issues to help with dependency coordination and tracking. fyi @kylebernhardy, since there's a good bit of related bits here
Audited this acceptance list against origin/main today and updated it in place. It had gone stale: every one of the 16 boxes was unchecked, body untouched since 2026-05-28, even though #848 merged 2026-06-02 and implemented most of the loop.
9 done, 2 partial, 5 not started. Verdicts came from reading resources/models/agentLoop.ts, resources/models/types.ts, and reference/models/tool-calling.md — not commit messages. Per-item notes are inline in the body.
Three things worth pulling out of the detail:
1. The five remaining items are mostly one item. Resource resolution, MCP resolution, transaction-log audit, and permissions-under-caller-identity all block on the same missing seam — dispatching to a Resource rather than a caller-supplied toolHandlers entry. That is #1740. Only the JSON Schema validator is independent. So this issue is closer to done than 9/16 suggests, and its critical path is #1740.
2. Two criteria have drifted from what shipped, and should be reconciled rather than left to look like gaps:
"Tool argument validation strict by default" — the shipped default is 'none'; 'strict' and 'lenient' throw 501 at loop entry. Either build the validator or amend the criterion. Harper uses Joi internally and passes JSON Schema through to backends, so adopting Ajv is its own decision.
"opts.conversation?: ConversationResource" — shipped as conversation?: ConversationAppender, a structural one-way sink deliberately not coupled to Add ConversationResource for agent memory and conversation state #511 so callers can bring their own store. I checked this box, since it works on both sync and streaming paths, but the shape is not what the line describes.
Also note opts.maxTokens in the original list shipped as maxToolTokens.
3. Permissions is the one to settle before more code lands. "Tool calls run with caller's identity" is unimplementable while handlers are caller-supplied — there is no Harper-enforced identity on a tool call today. Once #1740 makes model-chosen tool calls reach real Resources, that becomes a live authorization surface. Worth deciding in #1740's design note rather than discovering during implementation.
Add agent-loop orchestration /
toolMode: 'auto'toscope.modelsContext
The model-access API (#510) declares
toolMode: 'return' | 'auto'onGenerateOpts.'return'mode is trivial (backend returns tool-call requests; caller resolves them externally).'auto'mode is the in-process agent loop —scope.modelsresolves model-issued tool calls againstscope.resourcesand registered MCP tools, executes them, and re-invokes the backend until the model produces a terminal answer.That loop is substantial enough to deserve its own issue: tool resolution conventions, MCP integration surface, streaming behavior across tool-call boundaries, audit, error handling, and safety limits all need design work. Implementing it inside #510 would gate the model-access primitives behind 1-2 weeks of agent-loop design. This issue carries that work separately.
Once shipped,
toolMode: 'auto'is what unlocks Harper's compressed-stack agent story: tool calls resolve in-process against local Resources, audited via the standard transaction log. The loop itself is the same regardless of target; the invocation latency depends on the tool target — sub-millisecond for Resource-backed tools (in-process dispatch), normal HTTP for tools the unified registry resolves from external MCP servers. The compressed-stack win applies to the Resource-backed case.API contract from #510 that this fulfills
'return'mode lands with #510 itself.'auto'is type-declared by #510 but throws "not yet implemented" until this issue ships.Proposed semantics
The auto loop:
scope.models.generate(input, { tools, toolMode: 'auto', ...opts }).opts.toolArgValidation: 'strict' | 'lenient' | 'none', default'strict'— models hallucinate JSON shape often enough that this catches real bugs cheaply). Resolve each call againstscope.resources(or the unified MCP tool registry), execute, capture results.Safety limits
The orchestrator enforces three independent limits, any one of which trips a structured budget-exceeded abort with the partial conversation for debugging:
opts.maxToolIterations— bounds the count of tool-call rounds. Default 10.opts.maxTokens— bounds the cumulative prompt + completion + tool-output tokens across the whole loop. No default (uncapped). Lets dependent issues (e.g. Built-in Harper Agent Component harper-pro#676's built-in agent enforcing per-session token budgets) bound spend at the orchestrator rather than via post-hoc inspection ofanalytics.model_call.opts.maxCostUsd— bounds cumulative USD cost. Per-call cost is computed fromanalytics.model_call.gpu_ms/ token counts × the configured rate card (per-model). No default (uncapped). Resolves the per-session cost-cap ask from Built-in Harper Agent Component harper-pro#676.A single iteration with a large prompt + large tool result + large response can blow a budget on its own — count is not a proxy for spend.
Parallel tool calls
Both OpenAI and Anthropic backends emit multiple tool calls in a single assistant message. Default execution: parallel — matches model expectations, latency-optimal. Configurable via:
opts.toolParallelism: 'parallel' | 'serial'Within a parallel batch, mutations do not share a transaction boundary — each tool call is its own Resource dispatch and commits independently (see "Transaction scope per turn" below). Apps that need atomicity across multiple steps should wrap the multi-step intent in a single Resource method exposed as one tool, not rely on parallel-batch atomicity.
Tool result size cap
A tool returning 10MB inflates the next round-trip's prompt and can exceed the model's context window. The orchestrator enforces a per-tool-result size cap with a truncation marker:
opts.toolResultMaxBytes— default 64KB. Result exceeding the cap is truncated and a marker (e.g.[…truncated; full result <Nbytes> bytes]) is appended; the truncated form is what feeds back into the model.Cheap to add now; awkward to retrofit once apps depend on full results flowing through.
AbortSignal propagation across the loop
When the caller aborts a
generatemid-loop:AbortSignalon the dispatched Resource call, when supported).Important for long-running agent endpoints (e.g. SSE-served chat completions where the client disconnects).
Tool resolution
Once the native MCP server in #465 lands, the
scope.resourcesregistry and "registered MCP tools" collapse into one unified tool registry — the same set of tools is exposed to external MCP clients (via #618 Application profile) and to in-processscope.modelsorchestration.Sources of tools in
'auto'mode:@export-ed Resource producesget_*,search_*,create_*,update_*,delete_*tools (only the verbs the class implements, RBAC-filtered).mcpToolsdeclaration per #622. For non-verb methods or computed tools.Tool resolution is a lookup in the central tool registry from #615, which already handles RBAC-aware filtering.
scope.modelsconsults the same registry external MCP clients see.Tool name convention
Inherited from #618: tool names are
<verb>_<resource>(get_product,create_order, etc.) with sanitization for path separators. This orchestration uses the same naming so external MCP clients and in-processscope.modelscallers reason about the same set of tools.Auto-discovery via the unified registry
With #618 landed, every
@export-ed Resource is automatically a candidate tool — no explicit registration inGenerateOpts.toolsrequired. The caller can still narrow the exposed set per call (tools: ['search_product', 'get_product']) or pass an explicit list for tighter control. Default behavior in'auto'mode: the full RBAC-filtered registry the calling user can see.Streaming with tool calls
For
generateStreamwithtoolMode: 'auto':The existing
GenerateChunkshape (deltaContent?,deltaToolCalls?,finishReason?) supports this transition without changes.Audit and replication
Tool calls inherit Harper's standard machinery for free:
analytics.model_call(per Add unified model-access API (scope.models) #510). The model call records the tool-call count; individual tool calls show up in the transaction audit log alongside data changes.This is the compressed-stack advantage made concrete: an agent's actions are recorded the same way as a human user's actions, in the same store, queryable with the same primitives.
Transaction scope per turn
Each tool call commits its own transaction. If the agent calls
create_orderand thenupdate_inventoryand the second fails, the first is already committed. This is the right default for composability with non-agent callers (tools-as-Resources behave identically whether called by an agent or directly), but worth calling out explicitly:ConversationResource integration
Optional affordance for callers using
ConversationResource(#511): passopts.conversation?: ConversationResourceand the orchestrator appends each turn (user message, assistant message, tool calls, tool results) as it runs. Without it, conversation persistence is the caller's responsibility — they'd reimplement the same write pattern around every agent loop invocation.The orchestrator does not require a
ConversationResource. Stateless agent calls work fine. When provided, the integration:role: 'tool'turn.The contract: the orchestrator writes turns as it runs; the caller decides whether to provide a conversation at all.
Error handling per tool call
When a tool call fails, default behavior: append the error as the tool result (so the model can react and choose to retry, work around, or abort). Configurable via
opts.toolErrorMode:'recover'(default) — return error to model, continue loop'abort'— abort the loop, return error to callerReturn shape
On success,
scope.models.generate({ toolMode: 'auto' })returns the final assistant response — sameGenerateResultshape callers get from'return'mode, just withouttoolCalls(those were resolved internally). Low overhead; matches the "I just want the answer" common case.Opt in to the full sequence (assistant turns, tool calls, tool results, intermediate responses) via
opts.includeToolTrace: true. Returns the response plus an orderedtrace: ToolTraceEntry[]for inspection and debugging.On error or safety-abort (budget exceeded, abort signal, tool error with
'abort'mode), the trace is included regardless ofincludeToolTrace. Operators need that trace to diagnose; hiding debug info on failure is an antipattern.Acceptance
toolMode: 'auto'orchestration loop implemented end-to-end. (feat(models): toolMode 'auto' agent loop on scope.models.generate (#612) #848 — sync and streaming.)scope.resourcesworks for an@export-ed Resource. ([Models] toolMode 'auto' does not resolve tools against scope.resources (only caller-supplied toolHandlers) — #510 criterion, untracked #1740. v1 dispatchesopts.toolHandlers, a caller-supplied table.)opts.maxToolIterations,opts.maxTokens, andopts.maxCostUsdindependently enforced; any trip returns a structured budget-exceeded abort with the partial trace. (Partial.maxToolIterationsis a hard bound and enforced. The token/cost caps shipped asmaxToolTokens/maxCostUsd— note the rename frommaxTokensabove — and are best-effort on the sync path: they derive from backendusage, so a backend reporting none warns once and continues rather than silently no-opping. Both throw 501 ongenerateStreambecauseGenerateChunkcarries nousagein v1.maxCostUsdadditionally needs a rate card — the v1 cost function returns 0, so it only trips if a non-zero function is wired.BudgetExceededError.partialTraceis in place.)opts.toolParallelism: 'serial'runs serially. (Default'parallel'viaPromise.all; each handler is its own dispatch, no shared transaction.)opts.toolResultMaxBytes, default 64KB) with truncation marker. (Default65_536; the model sees the truncated form, the trace records the original size.)AbortControllercomposed with the caller's signal; the composed signal reaches both the innergenerateandToolHandlerContext.signal. Budget trips fireloopController.abort()so an in-flight LLM call cancels rather than being awaited.)toolParallelismintypes.tsand in the tool-calling doc.)finishReasonis treated as an internal signal — stripped from forwarded chunks and re-emitted as one terminal chunk after the assistant turn is persisted, so a consumer stopping on finish-reason cannot race past the conversation append.)opts.toolErrorMode; trace returned on error/abort even whenincludeToolTracewould otherwise be off. ('recover'default vs'abort';'abort'throws before the conversation sink sees the round's tool turns, so the store never holds an error turn the model never consumed.)opts.toolArgValidation. (Not started, and the shipped default inverts this line: v1 implements'none'as the default;'strict'and'lenient'are reserved on the type surface and throw 501 at loop entry. Needs a JSON Schema validator decision — Harper uses Joi internally and passes JSON Schema through to backends, so adopting Ajv is its own call. Either build it or amend this criterion to match the shipped default.)opts.conversation?: ConversationResourceintegration appends turns as the loop runs. (Shipped with a deliberate shape change: the option isconversation?: ConversationAppender— a structural one-way sink, intentionally not coupled to Add ConversationResource for agent memory and conversation state #511'sConversationResource, so callers can plug in their own store. Wired on both sync and streaming paths. Add ConversationResource for agent memory and conversation state #511 integration remains separate.)reference/models/tool-calling.mdcovers declaring tools, both modes, options, budgets and errors, conversation persistence, and reserved options. "How to register a Resource as a tool" cannot be written until [Models] toolMode 'auto' does not resolve tools against scope.resources (only caller-supplied toolHandlers) — #510 criterion, untracked #1740.)analytics.model_callrow). (The loop calls back throughmodels.generate(..., {toolMode: 'return'})per iteration so each backend round flows the single-shot path and writes its ownhdb_model_callsrow; the outer'auto'call stays out of the table.)Dependencies
Hard: model-access API (Add unified model-access API (scope.models) #510), specifically Phase 3 — needs a backend that natively supports tools (
openaibackend, thenanthropic/bedrockfollow).ollamadoes not currently expose tools in a portable way; if/when it does, this orchestration picks it up automatically.Hard: native MCP server family (#465 umbrella), specifically:
mcpToolsdeclaration (for non-verb methods)Previously this dependency was framed against the external
HarperFast/mcp-serveraddon; that path is superseded by the native MCP server (andHarperFast/mcp-serveris being archived per #623).Related
toolMode: 'auto'declared there.turns.tool_callsandturns.tool_resultsrecord what the orchestrator did; optional integration viaopts.conversationper the "ConversationResource integration" section above.Out of scope
opts.toolPermissionsto run the loop with a tighter permission set than the caller's). Tool calls inherit the caller's permissions in v1; tighter scoping is a v1.1 / follow-up consideration.maxToolIterationsalready catches infinite loops eventually. A "same tool, same args, N times consecutively" check would fail faster with a more actionable error — fast-follow if needed; not load-bearing for v1.🤖 Generated with Claude Code