Local quota protection and model stepping for Claude Code and Codex.
Usage Guard reads every short, weekly, and additional quota window the provider actually exposes, learns the observed quota burn rate, and makes a visible choice before a session runs into its protected reserve. It never invents a reset time when the provider has not supplied one.
Model stepping is on by default. Usage Guard recommends cheaper models for new work as quota tightens, preserves strong planning and review, and stops at your reserve. Every recommendation has a visible reason.
- Shows every quota bucket the provider exposes, including authoritative reset time when available and current observed pace.
- Protects separate reserves for short and weekly windows.
- Tracks Claude context pressure per session when Claude supplies it.
- Optionally surfaces current Claude request context, resident images, cache writes/reads, and directional weighted cost from a local pxpipe ledger.
- Keeps screenshots visually available for five later user turns, then replaces their repeated image payload with compact text memory on future Claude requests.
- Guides agents to reduce duplicate context and speculative concurrency.
- Recommends separate model/effort settings for the main thread and routine subagents, with recovery hysteresis and optional quiet hours.
- Offers an optional quality lock to retain the previous model-preserving behavior.
- Blocks new prompts and tool calls at the reserve and explains when work can resume. With model stepping active, it also saves a quota-only handover.
- Rechecks quota before Claude tool calls and after tool batches so an active desktop run can stop before another expensive model request.
- Uses Claude's native compaction at 40% by default and saves the provider's compact summary as a private project-keyed handoff for a fresh local session.
- Adds a Claude Code status line and Codex's built-in five-hour/weekly status segments.
- Shows the same deliberate-choice notice in CLI and desktop plugin sessions.
- Provides a local dashboard and MCP tools for agents.
Usage Guard cannot create unlimited included quota. If demand exceeds the provider allowance, uninterrupted operation requires waiting for reset or a separately enabled paid overflow path. The current release implements protected waiting, not paid overflow.
Node.js 22 or newer is required. Until the npm package is published, install directly from the public repository:
npm install --global https://github.com/agent-layer-mcp/usage-guard/releases/latest/download/usage-guard.tgz
usage-guard installThe current release is 0.5.1. Claude Code 2.1.284 or newer is required
for the default Opus/Sonnet 5.5 ladder. usage-guard doctor checks the CLI
version; an older CLI must be updated before applying those recommendations.
Claude Desktop has its own runtime and available tools.
The installer:
- Adds the Agent Layer marketplace and Usage Guard plugin to Claude Code and Codex.
- Configures Claude's custom status line.
- Adds Codex's built-in model, context, five-hour, and weekly status segments.
- Backs up affected settings and records the exact values needed for rollback.
- Records absolute Node and CLI paths and installs a
/bin/shbootstrap so plugin hooks and MCP startup do not depend on a TerminalPATH. - On macOS, builds a small native Usage Guard notification helper and installs
a local one-minute background monitor. It alerts when Claude usage escalates
to
WATCH,PROTECT, orQUEUE, with legacy notification paths retained only as fallbacks. - On macOS, installs Screenshot Memory as a localhost-only Claude request
gateway. It preserves any existing
ANTHROPIC_BASE_URLas its upstream and configures the Claude CLI route. Claude Desktop ignores that settings-file route and must be configured separately through its supported Third-Party Inference screen. - Configures Claude's supported proactive compaction variables to compact at 40% of the active model window. The installer preserves and restores the user's previous values exactly.
Claude and Codex require users to review and trust newly installed lifecycle hooks. Review the bundled hooks in plugins/usage-guard/hooks and approve them in the provider UI.
# Live terminal view. Supported local provider meters are refreshed automatically.
usage-guard status
# Refresh one provider or both explicitly.
usage-guard sync claude
usage-guard sync codex
usage-guard sync all
# Local dashboard.
usage-guard serve
# Verify providers, meters, and the privacy boundary.
usage-guard doctor
# Inspect or change policy.
usage-guard config
usage-guard config enforcement protect weeklyReservePercent 7
usage-guard config contextWatchPercent 70 contextProtectPercent 85
usage-guard config compactionHandoffEnabled false
usage-guard config qualityLock true
usage-guard config qualityLock false modelStepping true
# Choose a different native auto-compaction percentage during installation.
usage-guard install --auto-compact-percent 50The dashboard runs on http://127.0.0.1:4765 and binds only to localhost.
The default ladders use these exact IDs, verified against the providers' Claude model documentation and Codex model documentation:
| Rung | Claude model / effort | Codex model / effort |
|---|---|---|
| 1 | claude-opus-5-5 / high |
gpt-6-astra / high |
| 2 | claude-sonnet-5-5 / high |
gpt-6-sol / high |
| 3 | claude-sonnet-5-5 / medium |
gpt-6-sol / medium |
| 4, routine only | claude-haiku-4-5-20251001 / default |
gpt-6-luna / high |
Haiku's default is a Usage Guard label: omit the host's effort parameter and
clear inherited overrides. Haiku 4.5 does not support that parameter. The
stepped-down rungs reject max and xhigh. Fable and GPT-5.5 are not in the
default ladders. Haiku 5.5 is not preconfigured before release; a verified new
model can be added later through modelLadders.
Usable quota is remaining quota minus the protected reserve, in percentage
points of the whole allowance. Five-hour usable quota below 40%, 25%, and 12%
selects rungs 2, 3, and 4. Weekly usable quota below 30% selects rung 2; below
10% selects rung 4. The worse window wins. Rapid burn adds one rung, and quiet
hours can add another; the ladder stops at rung 4. A sustained-pace warning
(paceRatio >= 0.95) or a predicted reserve within rapidBurnWatchMinutes
counts as rapid burn. Context pressure alone does not step models down.
Each quota window retains its rung until usable quota reaches the triggering threshold plus the 10-point recovery margin. A fresh, provider-confirmed reset releases that window's latch without releasing another pressured window. Rapid-burn and quiet-hours adjustments end when those conditions end. Missing, stale, or expired meters withhold recommendations until refreshed; a fresh window at its reserve still stops work even if another meter is stale.
Planning, architecture, finished-work reviews, and work involving money, security, children's data, or production use the best rung allowed by quota, never below rung 2. Routine searches, test runs, reading, mechanical edits, and summaries go one rung cheaper first. General implementation can use rung 3; the floor is routine-only. Without a task description, status assumes the main thread is planning. Sensitive task classification overrides a supplied routine role, using the task description in memory only.
usage_guard_status and usage_guard_decision return recommendedModel,
recommendedEffort, routineModel, routineEffort, and a one-line
modelReason. Pass a role or transient task to obtain the right
recommendation before starting a different kind of work. The instructions
tell the agent to finish the current edit and verification, then switch between
tasks. In Claude Desktop, agents can use the exposed
mcp__ccd_session_mgmt__set_session_model and
mcp__ccd_session_mgmt__set_session_effort tools with their declared schemas.
Other hosts use an available session control or a new subagent with explicit
model and effort. A host without either capability must report the limitation.
Usage Guard emits guidance; it does not impersonate a host control or claim an
unconfirmed switch. Subagents share the same quota and cannot bypass a reserve
pause. Enabled stepping does not ask for confirmation on each change.
Changes are recorded in recentDecisions with timestamps, quota figures,
chosen rungs, and both model/effort pairs. These are recommendation changes,
marked applied: false, not proof a host changed its session. The audit tracks
the protected-main and routine roles consistently so querying a different role
does not invent overnight changes. A separate history retains steps even after
many status polls. overnightSummary summarizes the last 12 hours by default.
Every setting is configurable through usage_guard_configure, including
custom ladders, thresholds, reserves, hysteresis, role limits, rapid-burn
settings, quiet hours, and summary duration. See the complete default config
and example status. quietHours: null disables quiet
hours; { "start": "22:00", "end": "07:00" } uses the machine's local time.
An optional IANA timeZone supports a fixed zone, and extraRungs defaults to 1.
Setting qualityLock: true takes precedence over stepping and restores the
legacy provider decisions and hook behavior. Setting both switches to false
leaves model choice alone. Upgrading from the old fixed-lock release turns on
the new default once; explicitly re-enabling the lock thereafter persists.
| State | Behavior |
|---|---|
safe |
Work normally; avoid duplicate context. |
watch |
Reuse evidence, compact at safe boundaries, and bound parallel work. |
protect |
Use one active path, preserve strong planning/review, and stop speculative branches. |
queue |
Pause new prompts and tool calls at the reserve; save a quota-only handover when model stepping is active. |
missing / stale |
Refresh the meter; never block from uncertain data. |
Default reserves are 8% for windows up to six hours and 5% for longer windows. Change them locally with usage-guard config.
Claude context guidance defaults to watch at 70% and protect at 85%.
Context pressure never changes the selected model and never causes a hard block
by itself. At the protect threshold, Usage Guard asks the agent to finish its
current coherent step, update a durable handoff, and compact or start a fresh
task before another large phase.
Usage Guard configures Claude's documented
CLAUDE_CODE_AUTO_COMPACT_WINDOW and
CLAUDE_AUTOCOMPACT_PCT_OVERRIDE variables. The default is 40%. Claude still
performs the compaction itself with the active session's summarization
machinery. Usage Guard does not create a replacement summary; resumed work
follows the current model and effort policy.
After compaction, the supported PostCompact hook receives Claude's compact
summary. By default, Usage Guard stores that summary below
~/.usage-guard/handoffs/<project-key>/context-handoff.md with user-only
permissions. A different fresh local session for the same Git repository
receives the recent handoff through its supported SessionStart context. The
summary is not redacted or filtered and may contain prompt and response text,
source code, diffs, pasted secrets, or other sensitive content. Usage Guard does
not read or copy the raw transcript to create it. Disable future saving and
injection with:
usage-guard config compactionHandoffEnabled falseThis does not delete an existing handoff file; usage-guard reset removes saved
handoffs along with quota history and other local state.
Claude Desktop does not expose an official way for a plugin to silently create and submit a new Code session. Its deep link can only prefill a composer and always confirms a supplied folder. Native compaction in the same session is therefore the fully automatic continuity path.
Claude can send rate_limits.five_hour and rate_limits.seven_day data to the
Usage Guard status-line command. Eligible fields appear after a response and
individual windows may be absent. The payload also carries model, effort,
session, and context data. Usage Guard stores normalized quota observations and
local per-session aggregate context percentages, not transcript contents.
Claude Desktop's local Code surface runs Usage Guard's hooks and MCP server but
does not currently invoke the configured custom status-line command. On macOS,
Usage Guard therefore reads Claude Desktop's user-only aggregate plan history
before each prompt. It accepts only the observed fh and sd percentages and
their timestamps, ignores unknown fields, and never stores the cache's
organization identifier. This is an undocumented desktop cache, so malformed or
changed schemas fail open as a missing signal.
The Desktop cache does not include reset timestamps. Claude Code v2.1.80 and
later provides authoritative rate_limits.five_hour.resets_at and
rate_limits.seven_day.resets_at values to its supported status-line JSON.
Usage Guard records those values and can retain a still-future reset while the
Desktop cache refreshes the corresponding percentage. It discards that reset
after expiry or an observed usage rollover. Before Claude Code supplies an
authoritative sample, Usage Guard reports the reset as unknown; it never treats
an earlier aggregate drop as the anchor for a later window.
Usage Guard still projects minutes until the configured reserve from multiple
recent aggregate-burn horizons: WATCH within 90 minutes and PROTECT within
30 minutes by default.
Vision payloads are unusually expensive when the same screenshot remains in a long conversation. Screenshot Memory keeps each screenshot intact for the next five user turns. On later outbound Claude requests, it replaces the base64 image block with bounded text containing:
- the text from the user's original screenshot turn
- the first assistant response immediately following that turn
- the image type, approximate original size, and a short content hash
- a clear warning that the replacement is contextual memory, not a pixel-exact transcription
This transformation happens only in the request sent to Claude. It does not
edit the saved conversation or delete the original attachment. Reattach an
image whenever its pixels are needed again. Add [usage-guard:pin-images] or
#keep-screenshot to the screenshot turn to keep its image payload live.
The gateway binds to 127.0.0.1:47822. It does not call another model to create
the memory, and it stores replacement counts, byte totals, timestamps, retention
settings, and the configured upstream URL. On
install it chains to the user's prior Claude upstream and restores that exact
setting on uninstall.
Screenshot Memory applies to local Claude CLI requests after installation.
Claude Desktop explicitly does not use ANTHROPIC_BASE_URL or settings.json
for gateway routing. Its supported Third-Party Inference configuration uses a
separate gateway credential path rather than silently proxying the user's
normal claude.ai subscriber session. Usage Guard therefore does not enable or
claim Screenshot Memory for a normal Desktop subscription session. Users who
already operate a credentialed Desktop gateway can deliberately point that
configuration at http://127.0.0.1:47822 and verify it separately. See
Anthropic's
gateway documentation.
Codex does not offer an equivalent supported local request-routing surface, so Usage Guard does not claim to remove images from Codex context.
Screenshot Memory is a narrow local request gateway, not a hosted API service.
When the separately installed local pxpipe proxy has a fresh
~/.pxpipe/events.jsonl ledger, the status command, MCP status tool, and
escalation notification can surface request metadata that explains cost:
- baseline context tokens
- resident image count
- cache creation and cache-read tokens
- output tokens
- directional API-equivalent request weight
Cache writes are weighted at 1.25x for an explicitly reported five-minute
TTL and 2x for an explicitly reported one-hour TTL. If pxpipe does not identify
the TTL, Usage Guard shows the possible range instead of choosing the cheaper
multiplier. These ratios are useful diagnostics, not a claim about Anthropic's
private Max-plan quota formula. Provider-reported usage percentage and its
observed derivative remain the decision authority.
Usage Guard starts the local codex app-server, completes the documented initialization handshake, calls account/rateLimits/read, stores every returned limit bucket, and exits the child process. It does not read auth.json or provider tokens.
The meter separately reads config/read for the configured model and effort
when the hook did not supply a session-specific choice. A configured default is
not proof of a session override. Quota bucket labels are never treated as model
selection: a Spark-named bucket retains that name if the provider still returns
it, while buckets absent from a fresh authoritative response leave the current
view. Their history is retained for inspection.
Codex currently supports built-in quota status-line segments rather than a Claude-style arbitrary renderer. The plugin uses those native segments, while lifecycle notices and the local dashboard show richer decisions.
Codex's stable hook payload does not currently expose an exact context percentage to Usage Guard. Codex displays its own native context meter and automatically compacts. Usage Guard does not read the unstable transcript to manufacture an estimate.
- Codex Desktop Work/Codex sessions: local plugin, hooks, MCP, and native status segments.
- Codex CLI: plugin, hooks, MCP, and native status segments.
- Claude Desktop Code tab, local sessions: plugin, session/prompt/tool hooks, MCP,
aggregate five-hour/weekly cache ingestion, and background macOS
notifications. Sessions receive the configured model policy at start.
With quality lock enabled, entering
PROTECTretains the legacy tool approval. With stepping enabled, recommendations are delivered for the next task boundary without asking again for downgrade permission; reachingQUEUEdenies the tool and stops the run.PostToolBatchprovides a second stop before another model request. The compactStopfooter remains a best-effort host message and is not treated as the enforcement surface. Screenshot Memory requires the separate Desktop Third-Party Inference setup described above. - Claude Code CLI: plugin, hooks, MCP, and custom status line.
- Claude Desktop SSH and remote sessions: do not assume that the local desktop aggregate cache describes the remote account or that provider hooks are supported.
Run usage-guard doctor after installation. It verifies both CLIs, the
PATH-independent plugin bootstrap, absolute desktop runtime paths, status-line
configuration, native early compaction, plugin enabled state, Claude Desktop
aggregate-cache ingestion, meter freshness, and whether Claude has supplied a
context observation. It deliberately reports Claude Desktop gateway routing as
unverified instead of treating the Claude CLI environment as proof.
Claude Desktop displays exact context usage in its native UI, but its observed stable plugin hook payload and aggregate cache do not expose that percentage. Usage Guard does not read transcripts to manufacture it. Desktop users retain Claude's native context meter and automatic compaction; exact Usage Guard context guidance remains available where Claude's status-line payload supplies the value.
- Model stepping is visible, authorized by the enabled setting, and applied only between tasks.
- Planning, architecture, sensitive work, and final review never go below rung 2.
- Routine work steps down first; tests and verification remain required.
- Quality lock restores the previous model-preserving behavior.
- Stale or missing observations never justify a stop; fresh reserve evidence does.
- Every intervention names its reason and uses a reset time only when supplied.
Switching models can change cache economics, so the governor uses hysteresis and task boundaries to avoid repeated switches. API token prices are not a guarantee of subscription quota savings; those provider weights are private.
Usage Guard materially reduces surprise exhaustion when provider observations
are fresh. It cannot guarantee that an account never reaches a provider limit:
another device or application can consume the same quota, providers can omit a
window, and no plugin can interrupt a model generation already in flight.
Usage Guard runs a PreToolUse check before tool calls and a PostToolBatch
check before another model request. The background monitor also provides
escalation alerts. A reserve stop blocks new prompts and tool calls, including
deterministic tools inside the stopped run; independent local checks outside
that run are unaffected.
Usage Guard is local-only by default. Its quota governance and model-step records do not store prompt text, source code, diffs, or transcript bodies. The separate provider-supplied compact-summary handoff is enabled by default and is saved locally without redaction or filtering; it may contain exact prompt or response text, code, diffs, pasted credentials, and other sensitive content. Usage Guard does not read or copy the raw transcript to produce that handoff. Screenshot Memory inspects image payloads in memory but does not save them; it forwards authorization headers without persisting them. Quota governance does not directly read browser sessions or provider auth files.
It stores normalized quota observations, model/effort labels supplied by
lifecycle payloads, local session identifiers, aggregate context percentage,
policy decisions, model-step history, and local configuration in
~/.usage-guard/usage-guard.sqlite3.
Reserve handovers contain only quota/model metadata, never the task text,
prompts, source code, or tool arguments. With model stepping active
(modelStepping: true and qualityLock: false), the hook saves one locally
before stopping. This does not happen under the legacy quality lock. The
compact-summary handoff is separate and enabled by default.
On macOS its hook and one-minute background monitor may read Claude Desktop's
local aggregate plan history at
~/Library/Application Support/Claude/plan-usage-history.json. Usage Guard
stores only normalized percentages and timestamps from supported fields. It
does not store organization identifiers or unknown cache fields.
If pxpipe is installed separately, status surfaces may read its local request ledger for the allowlisted numeric diagnostics listed above. Usage Guard does not read Claude transcripts, does not return request bodies, and does not store pxpipe rows. If transcript support is ever added, duplicated streaming records must be grouped by request/message ID and the maximum usage values retained; naive line summation is explicitly prohibited.
Screenshot Memory transiently parses outbound Claude JSON requests in local memory to identify image blocks and their surrounding text. It forwards the request to the configured upstream and does not write prompt text, assistant text, image data, authorization headers, or transformed request bodies to disk. Its status file contains only counts, byte totals, timestamps, retention configuration, and the upstream URL.
See PRIVACY.md and SECURITY.md.
usage-guard uninstall
npm uninstall --global @agent-layer/usage-guardThe uninstall command removes both plugin integrations and restores the prior status-line settings without replacing unrelated current settings.
On macOS it also unloads the Usage Guard background monitor and removes the
native notification helper and Screenshot Memory gateway. Claude's prior
ANTHROPIC_BASE_URL is restored.
npm install
npm run check
npm test
USAGE_GUARD_HOME="$(mktemp -d)" node bin/usage-guard.mjs demoFor local marketplace testing:
codex plugin marketplace add "$PWD"
codex plugin add usage-guard@agent-layer
claude plugin marketplace add "$PWD"
claude plugin install usage-guard@agent-layer --scope userThis is an early public release. Claude and Codex can change quota shapes and plugin surfaces. The adapters intentionally treat windows as optional and arbitrary. Claude Desktop's aggregate cache is not a documented API and may change. A missing five-hour meter means the provider did not expose a supported reading in that snapshot, not that Usage Guard silently inferred it.
Built by Agent Layer. Licensed under MIT.