Skip to content

Add agent-loop orchestration / toolMode: 'auto' to scope.models #612

Description

@heskew

Add agent-loop orchestration / toolMode: 'auto' to scope.models

Context

The model-access API (#510) declares toolMode: 'return' | 'auto' on GenerateOpts. 'return' mode is trivial (backend returns tool-call requests; caller resolves them externally). 'auto' mode is the in-process agent loop — scope.models resolves model-issued tool calls against scope.resources and registered MCP tools, executes them, and re-invokes the backend until the model produces a terminal answer.

That loop is substantial enough to deserve its own issue: tool resolution conventions, MCP integration surface, streaming behavior across tool-call boundaries, audit, error handling, and safety limits all need design work. Implementing it inside #510 would gate the model-access primitives behind 1-2 weeks of agent-loop design. This issue carries that work separately.

Once shipped, toolMode: 'auto' is what unlocks Harper's compressed-stack agent story: tool calls resolve in-process against local Resources, audited via the standard transaction log. The loop itself is the same regardless of target; the invocation latency depends on the tool target — sub-millisecond for Resource-backed tools (in-process dispatch), normal HTTP for tools the unified registry resolves from external MCP servers. The compressed-stack win applies to the Resource-backed case.

API contract from #510 that this fulfills

type GenerateOpts = {
  tools?: ToolDef[];
  toolMode?: 'return' | 'auto';
  // ... other fields from HarperFast/harper#510
};

'return' mode lands with #510 itself. 'auto' is type-declared by #510 but throws "not yet implemented" until this issue ships.

Proposed semantics

The auto loop:

  1. Caller invokes scope.models.generate(input, { tools, toolMode: 'auto', ...opts }).
  2. Backend returns a result.
  3. If the result has no tool calls, return to caller. Done.
  4. If tool calls are present: validate each call's arguments against the tool's JSON Schema (opts.toolArgValidation: 'strict' | 'lenient' | 'none', default 'strict' — models hallucinate JSON shape often enough that this catches real bugs cheaply). Resolve each call against scope.resources (or the unified MCP tool registry), execute, capture results.
  5. Append the assistant's tool-call message and the tool results to the message list.
  6. Re-invoke the backend with the updated message list.
  7. Repeat from step 3 until terminal (no more tool calls) or any safety limit trips.

Safety limits

The orchestrator enforces three independent limits, any one of which trips a structured budget-exceeded abort with the partial conversation for debugging:

  • opts.maxToolIterations — bounds the count of tool-call rounds. Default 10.
  • opts.maxTokens — bounds the cumulative prompt + completion + tool-output tokens across the whole loop. No default (uncapped). Lets dependent issues (e.g. Built-in Harper Agent Component harper-pro#676's built-in agent enforcing per-session token budgets) bound spend at the orchestrator rather than via post-hoc inspection of analytics.model_call.
  • opts.maxCostUsd — bounds cumulative USD cost. Per-call cost is computed from analytics.model_call.gpu_ms / token counts × the configured rate card (per-model). No default (uncapped). Resolves the per-session cost-cap ask from Built-in Harper Agent Component harper-pro#676.

A single iteration with a large prompt + large tool result + large response can blow a budget on its own — count is not a proxy for spend.

Parallel tool calls

Both OpenAI and Anthropic backends emit multiple tool calls in a single assistant message. Default execution: parallel — matches model expectations, latency-optimal. Configurable via:

  • opts.toolParallelism: 'parallel' | 'serial'

Within a parallel batch, mutations do not share a transaction boundary — each tool call is its own Resource dispatch and commits independently (see "Transaction scope per turn" below). Apps that need atomicity across multiple steps should wrap the multi-step intent in a single Resource method exposed as one tool, not rely on parallel-batch atomicity.

Tool result size cap

A tool returning 10MB inflates the next round-trip's prompt and can exceed the model's context window. The orchestrator enforces a per-tool-result size cap with a truncation marker:

  • opts.toolResultMaxBytes — default 64KB. Result exceeding the cap is truncated and a marker (e.g. […truncated; full result <Nbytes> bytes]) is appended; the truncated form is what feeds back into the model.

Cheap to add now; awkward to retrofit once apps depend on full results flowing through.

AbortSignal propagation across the loop

When the caller aborts a generate mid-loop:

  1. Cancel the upstream LLM call in flight.
  2. Propagate the abort into in-flight tool executions via the tool's invocation context (AbortSignal on the dispatched Resource call, when supported).
  3. Already-committed writes from completed tool calls stay committed — no implicit rollback. The orchestrator returns a structured abort error.

Important for long-running agent endpoints (e.g. SSE-served chat completions where the client disconnects).

Tool resolution

Once the native MCP server in #465 lands, the scope.resources registry and "registered MCP tools" collapse into one unified tool registry — the same set of tools is exposed to external MCP clients (via #618 Application profile) and to in-process scope.models orchestration.

Sources of tools in 'auto' mode:

  1. Auto-generated from the Resources registry via #618. Each @export-ed Resource produces get_*, search_*, create_*, update_*, delete_* tools (only the verbs the class implements, RBAC-filtered).
  2. Custom Resource methods opted in via static mcpTools declaration per #622. For non-verb methods or computed tools.

Tool resolution is a lookup in the central tool registry from #615, which already handles RBAC-aware filtering. scope.models consults the same registry external MCP clients see.

Tool name convention

Inherited from #618: tool names are <verb>_<resource> (get_product, create_order, etc.) with sanitization for path separators. This orchestration uses the same naming so external MCP clients and in-process scope.models callers reason about the same set of tools.

Auto-discovery via the unified registry

With #618 landed, every @export-ed Resource is automatically a candidate tool — no explicit registration in GenerateOpts.tools required. The caller can still narrow the exposed set per call (tools: ['search_product', 'get_product']) or pass an explicit list for tighter control. Default behavior in 'auto' mode: the full RBAC-filtered registry the calling user can see.

Streaming with tool calls

For generateStream with toolMode: 'auto':

  • Yields content deltas as they arrive from the backend.
  • When the model emits a tool-call delta, the iterator yields that delta (signaling the consumer; useful for UI "agent is calling X" indicators).
  • The orchestrator suspends the upstream stream, executes the tool call(s), feeds results back to the backend.
  • The new stream from the backend resumes, yielding more deltas.
  • Loop until terminal.

The existing GenerateChunk shape (deltaContent?, deltaToolCalls?, finishReason?) supports this transition without changes.

Audit and replication

Tool calls inherit Harper's standard machinery for free:

  • Each tool call goes through the Resource dispatch — auth chain runs (tool calls have the caller's permissions), the call is logged in the transaction log, mutations replicate.
  • The model call itself is logged in analytics.model_call (per Add unified model-access API (scope.models) #510). The model call records the tool-call count; individual tool calls show up in the transaction audit log alongside data changes.
  • Querying "what did this agent actually do" is a regular Harper query against the audit log filtered by request / conversation / tenant.

This is the compressed-stack advantage made concrete: an agent's actions are recorded the same way as a human user's actions, in the same store, queryable with the same primitives.

Transaction scope per turn

Each tool call commits its own transaction. If the agent calls create_order and then update_inventory and the second fails, the first is already committed. This is the right default for composability with non-agent callers (tools-as-Resources behave identically whether called by an agent or directly), but worth calling out explicitly:

  • Multi-step atomic intent should be wrapped in a single Resource method (exposed as one tool), not relied on across multiple tool calls.
  • The orchestrator does not provide automatic rollback or saga semantics across the loop.

ConversationResource integration

Optional affordance for callers using ConversationResource (#511): pass opts.conversation?: ConversationResource and the orchestrator appends each turn (user message, assistant message, tool calls, tool results) as it runs. Without it, conversation persistence is the caller's responsibility — they'd reimplement the same write pattern around every agent loop invocation.

The orchestrator does not require a ConversationResource. Stateless agent calls work fine. When provided, the integration:

  • Appends the input as a user/system turn before the loop starts.
  • For each tool call: appends the assistant's tool-call turn and the tool result(s) as a role: 'tool' turn.
  • Appends the final assistant response as a turn before returning.
  • Uses Add ConversationResource for agent memory and conversation state #511's streaming-append protocol (open → chunk → commit) on the streaming path.

The contract: the orchestrator writes turns as it runs; the caller decides whether to provide a conversation at all.

Error handling per tool call

When a tool call fails, default behavior: append the error as the tool result (so the model can react and choose to retry, work around, or abort). Configurable via opts.toolErrorMode:

  • 'recover' (default) — return error to model, continue loop
  • 'abort' — abort the loop, return error to caller

Return shape

On success, scope.models.generate({ toolMode: 'auto' }) returns the final assistant response — same GenerateResult shape callers get from 'return' mode, just without toolCalls (those were resolved internally). Low overhead; matches the "I just want the answer" common case.

Opt in to the full sequence (assistant turns, tool calls, tool results, intermediate responses) via opts.includeToolTrace: true. Returns the response plus an ordered trace: ToolTraceEntry[] for inspection and debugging.

On error or safety-abort (budget exceeded, abort signal, tool error with 'abort' mode), the trace is included regardless of includeToolTrace. Operators need that trace to diagnose; hiding debug info on failure is an antipattern.

Acceptance

Audited against origin/main 2026-08-06. This list had gone stale — every box was unchecked despite #848 (merged 2026-06-02) implementing most of the loop. Verdicts below are from reading resources/models/agentLoop.ts, resources/models/types.ts, and HarperFast/documentation reference/models/tool-calling.md, not from commit messages. 9 done, 2 partial, 5 not started — the two partials (budgets, docs) are left unchecked.

The five not-started items are not five independent pieces of work: four of them (resource resolution, MCP resolution, transaction-log audit, permissions) all block on the same missing seam — dispatching to a Resource instead of a caller-supplied handler — which is #1740. The fifth is a JSON Schema validator decision.

  • toolMode: 'auto' orchestration loop implemented end-to-end. (feat(models): toolMode 'auto' agent loop on scope.models.generate (#612) #848 — sync and streaming.)
  • Tool resolution against scope.resources works for an @export-ed Resource. ([Models] toolMode 'auto' does not resolve tools against scope.resources (only caller-supplied toolHandlers) — #510 criterion, untracked #1740. v1 dispatches opts.toolHandlers, a caller-supplied table.)
  • Tool resolution against MCP tools registered via the native MCP server family ([MCP] Tool registry + class-level introspection + RBAC-aware tools/list filtering #615 / [MCP] Application profile: tool generation over Resources registry #618 / [MCP] Custom Resource opt-in via static mcpTools declaration #622) works. (The MCP side landed; this needs [Models] toolMode 'auto' does not resolve tools against scope.resources (only caller-supplied toolHandlers) — #510 criterion, untracked #1740's seam to resolve through.)
  • opts.maxToolIterations, opts.maxTokens, and opts.maxCostUsd independently enforced; any trip returns a structured budget-exceeded abort with the partial trace. (Partial. maxToolIterations is a hard bound and enforced. The token/cost caps shipped as maxToolTokens / maxCostUsd — note the rename from maxTokens above — and are best-effort on the sync path: they derive from backend usage, so a backend reporting none warns once and continues rather than silently no-opping. Both throw 501 on generateStream because GenerateChunk carries no usage in v1. maxCostUsd additionally needs a rate card — the v1 cost function returns 0, so it only trips if a non-zero function is wired. BudgetExceededError.partialTrace is in place.)
  • Parallel tool-call execution by default; opts.toolParallelism: 'serial' runs serially. (Default 'parallel' via Promise.all; each handler is its own dispatch, no shared transaction.)
  • Per-tool-result size cap (opts.toolResultMaxBytes, default 64KB) with truncation marker. (Default 65_536; the model sees the truncated form, the trace records the original size.)
  • AbortSignal propagation: caller-side abort cancels upstream LLM call AND in-flight tool executions; already-committed writes stay committed. (Loop-level AbortController composed with the caller's signal; the composed signal reaches both the inner generate and ToolHandlerContext.signal. Budget trips fire loopController.abort() so an in-flight LLM call cancels rather than being awaited.)
  • Transaction scope per turn explicitly documented — each tool call commits independently. (Stated on toolParallelism in types.ts and in the tool-calling doc.)
  • Streaming + tool calls: chunks yield correctly through tool-execution transitions. (finishReason is treated as an internal signal — stripped from forwarded chunks and re-emitted as one terminal chunk after the assistant turn is persisted, so a consumer stopping on finish-reason cannot race past the conversation append.)
  • Tool calls audited in the transaction log alongside data writes (inherited from Resource dispatch). (Blocked on [Models] toolMode 'auto' does not resolve tools against scope.resources (only caller-supplied toolHandlers) — #510 criterion, untracked #1740 — the parenthetical says it: the audit is inherited from Resource dispatch, and there is no Resource dispatch yet. Caller-supplied handlers are opaque to the transaction log.)
  • Permissions enforced — tool calls run with caller's identity. (Blocked on [Models] toolMode 'auto' does not resolve tools against scope.resources (only caller-supplied toolHandlers) — #510 criterion, untracked #1740, and the load-bearing item in this set: with caller-supplied handlers there is no Harper-enforced identity on a tool call. Settle this before resource resolution ships, not after.)
  • Error handling per opts.toolErrorMode; trace returned on error/abort even when includeToolTrace would otherwise be off. ('recover' default vs 'abort'; 'abort' throws before the conversation sink sees the round's tool turns, so the store never holds an error turn the model never consumed.)
  • Tool argument validation strict by default; configurable per opts.toolArgValidation. (Not started, and the shipped default inverts this line: v1 implements 'none' as the default; 'strict' and 'lenient' are reserved on the type surface and throw 501 at loop entry. Needs a JSON Schema validator decision — Harper uses Joi internally and passes JSON Schema through to backends, so adopting Ajv is its own call. Either build it or amend this criterion to match the shipped default.)
  • Optional opts.conversation?: ConversationResource integration appends turns as the loop runs. (Shipped with a deliberate shape change: the option is conversation?: ConversationAppender — a structural one-way sink, intentionally not coupled to Add ConversationResource for agent memory and conversation state #511's ConversationResource, so callers can plug in their own store. Wired on both sync and streaming paths. Add ConversationResource for agent memory and conversation state #511 integration remains separate.)
  • Documented: building an agent with auto mode; how to register a Resource as a tool; tool error handling; budget enforcement; abort semantics; conversation integration. (Partial. reference/models/tool-calling.md covers declaring tools, both modes, options, budgets and errors, conversation persistence, and reserved options. "How to register a Resource as a tool" cannot be written until [Models] toolMode 'auto' does not resolve tools against scope.resources (only caller-supplied toolHandlers) — #510 criterion, untracked #1740.)
  • Per-tenant accounting on the model calls is preserved through the loop (each round-trip writes its own analytics.model_call row). (The loop calls back through models.generate(..., {toolMode: 'return'}) per iteration so each backend round flows the single-shot path and writes its own hdb_model_calls row; the outer 'auto' call stays out of the table.)

Dependencies

  • Hard: model-access API (Add unified model-access API (scope.models) #510), specifically Phase 3 — needs a backend that natively supports tools (openai backend, then anthropic / bedrock follow). ollama does not currently expose tools in a portable way; if/when it does, this orchestration picks it up automatically.

  • Hard: native MCP server family (#465 umbrella), specifically:

    • #615 — Tool registry + class-level introspection + RBAC-aware filtering (the registry this orchestration consults)
    • #618 — Application profile (auto-generates per-Resource tools the orchestration resolves against)
    • #622 — Custom Resource opt-in via static mcpTools declaration (for non-verb methods)

    Previously this dependency was framed against the external HarperFast/mcp-server addon; that path is superseded by the native MCP server (and HarperFast/mcp-server is being archived per #623).

Related

Out of scope

  • Tool catalog / discovery UX in Studio.
  • Cross-tenant tool invocation.
  • A full multi-agent orchestration layer (this issue does single-agent tool-call loops; multi-agent coordination is its own design space).
  • Per-model-call permission scoping (opts.toolPermissions to run the loop with a tighter permission set than the caller's). Tool calls inherit the caller's permissions in v1; tighter scoping is a v1.1 / follow-up consideration.
  • Structural cycle detection. maxToolIterations already catches infinite loops eventually. A "same tool, same args, N times consecutively" check would fail faster with a more actionable error — fast-follow if needed; not load-bearing for v1.

🤖 Generated with Claude Code

Activity

  1. kriszyp commented on May 20, 2026

    @kriszyp
    Member

    This looks great! Let's move forward with this!
    Notes from Claude conversation:
    Foundation looks solid — the loop semantics, RBAC-via-Resource-dispatch, and unified-registry anchor are all the right calls. A few minor gaps / adjustments worth folding in (none are architectural; mostly defaults and safety limits that are cheap to add now and awkward to retrofit):

    1. Cost cap, not just iteration cap. maxToolIterations bounds count, not spend. One iteration with a large prompt + large tool result + large response can blow a budget on its own. Add opts.maxTokens and/or opts.maxCostUsd per generate call, with a structured budget-exceeded abort error. This also lets dependent issues (e.g. the built-in agent in Built-in Harper Agent Component harper-pro#676) enforce session budgets through the orchestrator rather than externally polling analytics.model_call.

    2. Parallel tool calls — explicit semantics. Both OpenAI and Anthropic backends emit multiple tool calls in a single assistant message. The proposal doesn't specify execution order. Recommend parallel by default (matches model expectation, latency-optimal) with opts.toolParallelism: 'parallel' | 'serial' opt-out, and document that mutations within a parallel batch don't share a transaction boundary spanning the batch.

    3. Transaction scope per turn — explicit decision. Each tool call goes through Resource dispatch and commits its own transaction. If the agent does create_order then update_inventory and the second fails, the first is already committed. That's almost certainly the right default (composability with non-agent callers), but worth calling out explicitly in the acceptance criteria so app authors know to wrap multi-step atomic intent in a single Resource method (called as one tool) when they need atomicity.

    4. AbortSignal propagation. Spec what happens when the caller aborts a generate mid-loop: (a) cancel the upstream LLM call, (b) propagate abort into in-flight tool executions via context.abortSignal, (c) already-committed writes from completed tool calls stay committed (no implicit rollback). Important for long-running agent endpoints.

    5. Tool result size cap. A tool returning 10MB inflates the next round-trip's prompt and can exceed the model's context window. Recommend a default per-tool-result size cap (e.g. ~64KB) with truncation marker, configurable via opts.toolResultMaxBytes. Cheap to add now; painful to retrofit once apps depend on full responses flowing through.

    6. Tool argument validation default (open decision Add a CODEOWNERS file to set PR reviewers #5). Recommend strict by default — models hallucinate JSON shape often enough that this catches real bugs cheaply. opts.toolArgValidation: 'strict' | 'lenient' | 'none' for opt-out.

    7. includeToolTrace default (open decision Default to attempting to serve index.html for a path #4). Recommend: final response only on success (as proposed), but always include the trace on error / safety-abort. Defaults that hide debug info on failure are an antipattern — operators need that trace to diagnose.

    8. ConversationResource integration boundary. "Pairs with Add ConversationResource for agent memory and conversation state #511" is mentioned, but the contract isn't specified: does the orchestrator write turns to a ConversationResource as it iterates, or is conversation persistence the caller's responsibility? If it's caller-side, every app reimplements the same persistence loop. Recommend a small affordance: opts.conversation?: ConversationResource that the orchestrator appends to as it runs. Out of scope to require it, but worth supporting.

    9. Latency claim precision. The "sub-millisecond invocation latency vs. MCP-over-HTTP" framing holds for Resource-backed tools (in-process dispatch) but not for tools the unified registry resolves from external MCP servers. Worth being precise — the loop is the same; the tool target determines latency. Otherwise a benchmark against an MCP-backed tool will appear to contradict the claim.

    10. (Minor) Cycle detection. maxToolIterations catches loops eventually; a structural cycle check ("same tool, same args, N times in a row") fails faster with a more actionable error. Optional, low priority.

    Most of these can land as additions to the existing "Open decisions" / "Acceptance" sections rather than new design work. Happy to PR a body edit if useful.

    — Claude

  2. heskew commented on May 20, 2026

    @heskew
    ContributorAuthor

    Thanks — folded all 10 into the body. Summary of what landed:

    • Safety limits section covers maxToolIterations + maxTokens + maxCostUsd independently; maxCostUsd derives per-call cost from analytics.model_call token counts × per-model rate card. Closes the per-session budget-enforcement ask from Built-in Harper Agent Component harper-pro#676.
    • Parallel tool calls subsection: parallel by default, opts.toolParallelism: 'serial' opt-out, with the explicit note that a parallel batch doesn't share a transaction boundary.
    • Tool result size cap subsection: opts.toolResultMaxBytes default 64KB with truncation marker.
    • AbortSignal propagation across the loop: explicit semantics for upstream cancel + in-flight tool cancel + committed-writes-stay-committed.
    • Transaction scope per turn explicit note under Audit: each tool call commits independently; atomic intent belongs in a single Resource method.
    • ConversationResource integration section: optional opts.conversation — orchestrator appends turns when provided; stateless calls still work.
    • Latency claim tightened — sub-millisecond applies to Resource-backed tools (in-process dispatch), not to tools the registry resolves from external MCP servers. The loop is the same; tool target drives latency.
    • toolArgValidation folded into step 4 of the loop semantics: strict by default; 'strict' | 'lenient' | 'none' opt-out.
    • includeToolTrace folded into a new "Return shape" section: final on success; full trace on opt-in; trace always included on error or safety-abort.
    • Cycle detection + per-model-call permission scoping moved to Out of scope as fast-follow / v1.1.

    Acceptance checklist updated to match. Open-decisions section dropped — decisions live in the body now, future work in Out of scope.


    🤖 Posted by Claude on Nathan's behalf

  3. self-assigned this
    on May 28, 2026
  4. heskew commented on May 28, 2026

    @heskew
    ContributorAuthor

    @kriszyp I broke this out into sub-issues to help with dependency coordination and tracking. fyi @kylebernhardy, since there's a good bit of related bits here

  5. heskew commented on Aug 7, 2026

    @heskew
    ContributorAuthor

    Audited this acceptance list against origin/main today and updated it in place. It had gone stale: every one of the 16 boxes was unchecked, body untouched since 2026-05-28, even though #848 merged 2026-06-02 and implemented most of the loop.

    9 done, 2 partial, 5 not started. Verdicts came from reading resources/models/agentLoop.ts, resources/models/types.ts, and reference/models/tool-calling.md — not commit messages. Per-item notes are inline in the body.

    Three things worth pulling out of the detail:

    1. The five remaining items are mostly one item. Resource resolution, MCP resolution, transaction-log audit, and permissions-under-caller-identity all block on the same missing seam — dispatching to a Resource rather than a caller-supplied toolHandlers entry. That is #1740. Only the JSON Schema validator is independent. So this issue is closer to done than 9/16 suggests, and its critical path is #1740.

    2. Two criteria have drifted from what shipped, and should be reconciled rather than left to look like gaps:

    • "Tool argument validation strict by default" — the shipped default is 'none'; 'strict' and 'lenient' throw 501 at loop entry. Either build the validator or amend the criterion. Harper uses Joi internally and passes JSON Schema through to backends, so adopting Ajv is its own decision.
    • "opts.conversation?: ConversationResource" — shipped as conversation?: ConversationAppender, a structural one-way sink deliberately not coupled to Add ConversationResource for agent memory and conversation state #511 so callers can bring their own store. I checked this box, since it works on both sync and streaming paths, but the shape is not what the line describes.

    Also note opts.maxTokens in the original list shipped as maxToolTokens.

    3. Permissions is the one to settle before more code lands. "Tool calls run with caller's identity" is unimplementable while handlers are caller-supplied — there is no Harper-enforced identity on a tool call today. Once #1740 makes model-chosen tool calls reach real Resources, that becomes a live authorization surface. Worth deciding in #1740's design note rather than discovering during implementation.


    🤖 Posted by Claude on Nathan's behalf

  6. added theissue type on Sep 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Fields

Priority

P2

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions