Skip to content

feat: policy-based rate limiting and cost-aware token budgeting #12

Description

Summary

Token budgets exist (daily + per-request) but there is no policy-based rate limiting on tool calls, concurrent requests, or cost-weighted budgeting.

Details

Gap 1 — No rate limiting per tool type (routes.rs:399-413)
Token budgets track raw token count, but there is no:

  • "Max 100 shell tool calls per day"
  • "Max 10 concurrent inference requests"
  • Rate limiting at the AGT policy level (despite the YAML policy format supporting rate_limit type)

Gap 2 — No cost-aware budgeting (budget.rs)
Token budget treats all tokens equally. GPT-4 tokens cost ~5x more than GPT-4o-mini tokens. A policy should support cost-based limits ("max $10/day" not just "max 1M tokens/day").

Gap 3 — Post-response token recording (routes.rs:465-478)
Token usage is recorded after the response. If the response stream fails midway, usage is not recorded, potentially allowing budget overruns.

Proposed Fix

AGT's rate limiter + CostGuard patterns:

// Rate limiting via AGT policy
let rate_decision = policy.evaluate_rate_limit("tool:shell", sandbox)?;
if !rate_limiter.allow(sandbox, rate_decision.max_calls) {
    return too_many_requests();
}

// Cost-aware budget
let cost = estimate_cost(model, estimated_tokens);
if budget.would_exceed(sandbox, cost) {
    return budget_exceeded();
}

References

  • routes.rs lines 399-413
  • budget.rs (entire file)
  • AGT: CostGuard, SRE error budgets

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions