Skip to content

RFC 0002: elastic execution — worker pools, provider governor, tiered models - #260

Open
pd-admin-jc wants to merge 3 commits into
mainfrom
spec/elastic-execution
Open

pd-admin-jc wants to merge 3 commits into
mainfrom
spec/elastic-execution

Conversation

@pd-admin-jc

Copy link
Copy Markdown

RFC 0002: elastic execution. One agent scales across pool members while its work queues, a provider governor paces every model call per deployment, and tiered models let cheap turns run on a cheaper model.

Design goals:

  • No developer configuration. Limits are learned from provider headers; pool sizes and lanes have working defaults.
  • Both profiles benefit. AutoAgent and CustomAgent get the same behaviour.

Implementation PRs:

  • feat/provider-governor: the provider governor (PR 1).
  • feat/elastic-pools: elastic pools, stacked on the governor (PR 2).
  • PR 3, tiered models, and PR 4, adoption and benchmark, are still to come.

Copilot AI balanced review requested due to automatic review settings October 11, 2026 09:06

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The RFC contains unresolved scaling, quota-accounting, concurrency, and model-routing design flaws.

11 open findings
What changed in this PR

Defines RFC 0002 for elastic execution in JarvisCore 2.2.

Changes:

  • Proposes elastic worker pools and distributed backlog scaling.
  • Designs provider-aware rate governance and priority lanes.
  • Introduces per-turn model tiers with acceptance criteria.
File Description
rfcs/​0002-elastic-execution.md Documents the elastic execution architecture and delivery plan.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.


- **Durable leased work.** `RedisContextStore.claim_step` leases a step atomically to
one execution identity. Leases expire and are recovered, and committing requires
holding the lease. Exactly-once execution and resume after a crash depend on this.
Each process runs one scaler coroutine per pool, once a second:

```
desired = clamp(ceil(backlog / per_worker), min, max)

```
desired = clamp(ceil(backlog / per_worker), min, max)
desired = min(desired, governor.headroom_workers(pool.model_tier))
Comment on lines +112 to +114
- **Across processes:** members register in `pool_members:{capability}` with a
heartbeat TTL. The global `max` is enforced there, so two API replicas don't
each start 32.
Comment on lines +138 to +140
1. **Provider metadata, when credentials allow.** For Azure, the deployment's
`sku.capacity` from Resource Manager is the TPM in thousands. RPM follows the
documented ratio of 6 RPM per 1,000 TPM for standard deployments.

### 5.2 Admission

Every call reserves capacity before it is dispatched:
cost = estimate(prompt_tokens) + max_tokens # what the provider charges at admission
reserve(key, requests=1, tokens=cost, lane) -> admitted | wait(seconds)
... call ...
settle(key, reservation, actual_total_tokens) # returns the unused estimate
Comment on lines +189 to +192
Lanes are `interactive` (chat turns), `decision` (final work products, verdicts)
and `bulk` (research loops). Bulk is admitted only while headroom stays above a
reserve, by default 15 % of TPM. A user's chat question during a 196-claim run is
therefore not queued behind it.
Comment on lines +211 to +214
- **Loop turns** use the fast tier: choosing tools, issuing searches, reading,
extracting quotes.
- **The final turn** uses the strong tier: the turn that emits `DONE/RESULT` for a
work-product contract.
|---|---|---|
| Governor (§5) | automatic | automatic: every profile calls the same `self.llm` client |
| Pools (§4) | `mesh.add(..., pool=Pool(...))` | the same, for agents that keep no per-instance state |
| Tiered turns (§6) | automatic: the kernel knows which turn is final | per call: `self.llm.generate(..., tier="fast" \| "strong")`; omit it for the default |

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants