Repository navigation
[v0.5.0] benchmark tokens — Claude token-count integration after client rewrap (live-phase-gated) #89
Description
Activity
- addedenhancementNew feature or requestNew feature or requestpriority:P2Medium priorityMedium priority
on Jul 8, 2026 Pre-flight amendments (2026-07-08, maintainer-approved — sign-off mirrored on #63)
Two binding changes to this issue's draft criteria:
-
Dependency rule TIGHTENED (supersedes the body's "optional extra" language; Codex round-1 finding 2): the worker may NOT add any dependency — no
anthropicSDK, no optional extra, no dev/benchmark group.pyproject.tomlanduv.lockmust be absent from the diff entirely; the PR body must includegit diff --stat main...HEADproving it. Implement the count-tokens call via the standard library or already-installed dependencies, strictly behind the live-phase guard. If an SDK appears genuinely required: stop and comment (pipeline §8); any dependency addition is a maintainer-owned commit followed by re-pre-flight. -
Named merged-test sanction (D6b): registering
"anthropic": "ANTHROPIC_API_KEY"inPROVIDER_API_KEY_ENV(benchmarks/adapters/guard.py:28-31) invalidates the merged test asserting"anthropic"is an unknown provider (tests/benchmarks/test_adapters.py). The sanction covers exactly: re-pointing that one assertion at a different unknown-provider string. Nothing else — any other merged-test change remains stop-and-comment.
Refs #63
-
Agent context — #89 [v0.5.0] benchmark tokens — Claude token-count integration after client rewrap (live-phase-gated)
Status: DRAFT for dispatch (verified against main @ ab9e615, 2026-07-08)
Auditable copy: this content is mirrored as a comment on issue #89 (.planning/is gitignored — the issue comment is the copy pre-flight verifies and the dispatch packet links).1. Roadmap excerpt (why this exists)
- Locked decision (PLAN Amendment 1, restated in [v0.5.0] benchmark tokens — Claude token-count integration after client rewrap (live-phase-gated) #89): the Anthropic count-tokens API is the counting mechanism, confined to the maintainer-run live phase. Exact counts back headline claims; the server itself gains zero runtime network access and zero new dependencies.
- Methodology "Token Measurement" steps 1–3 (
docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md): count the envelope after client-side rewrap; where a client cannot expose its exact wrapped envelope, mark the recordapproximation: true— approximations are never headline-eligible. - BINDING pre-flight amendment (issue [v0.5.0] benchmark tokens — Claude token-count integration after client rewrap (live-phase-gated) #89 comment, 2026-07-08 — supersedes the body's "optional extra" language): the worker may NOT add any dependency — no
anthropicSDK, no optional extra, no dev/benchmark group.pyproject.tomlanduv.lockmust be absent from the diff entirely. Implement the count-tokens call via the standard library or already-installed dependencies, strictly behind the live-phase guard. - The same comment grants the D6b merged-test sanction — scope in section 4. Both amendments are binding on this dispatch.
2. Code touch-points (verify against merged main before starting)
benchmarks/adapters/guard.py:54-81—require_live_environment(provider): the count-tokens call MUST sit behind this guard (flag + key, both required).benchmarks/adapters/guard.py:25—LIVE_PROVIDERS_ENABLED_ENV = "BENCHMARK_LIVE_PROVIDERS_ENABLED": the explicit opt-in flag; reuse it, do not invent a second flag.benchmarks/adapters/guard.py:28-31—PROVIDER_API_KEY_ENV: register"anthropic": "ANTHROPIC_API_KEY"here.benchmarks/adapters/guard.py:43-51—LiveExecutionNotImplementedError: the existing pattern for "guard passed but structurally incapable"; follow it for any path CI could reach.benchmarks/adapters/base.py:49-64—TokenMetadata: provider-native counts are diagnostic only and must never be reported underMETHODOLOGY_TOKEN_LABEL.benchmarks/model_matrix.py:33—METHODOLOGY_TOKEN_LABEL = "Claude Tokens (Normalized Payload)": the only label for methodology counts; use the exported constant, never a string literal.benchmarks/runner.py:250-260— token_record dict;client_wrapped_tokens/raw_payload_tokensplaceholders areNoneat :257-258. Fill both (raw payload counted separately, diagnostic). Update the stalenotesstring at :259 ("out of scope for issue [v0.5.0] benchmark runner — add reproducible CLI and artifact layout #72").benchmarks/report.py— CellRecord token fields :102-103, parsed :223-224;_basic_statstoken aggregation :355-379; familyheadline_eligiblerendering :480-481. Additive change: excludeapproximation: truerecords from headline-eligible tables.- Count-tokens HTTP call: stdlib (
urllib.request) or already-installed deps ONLY, invoked strictly afterrequire_live_environment("anthropic")passes.
3. Test patterns to follow
- Guard env monkeypatch pattern:
tests/benchmarks/test_adapters.py:106-148—monkeypatch.delenv/setenvonLIVE_PROVIDERS_ENABLED_ENV+PROVIDER_API_KEY_ENV[...], assertLiveProviderDisabledErrorwithmatch=. - Use a fake counter for ALL tests; CI and unit tests never call the real API. Prove structurally that no test path reaches the network.
- Cover, additively under
tests/benchmarks/: fake-counter record filling; approximation marking + headline exclusion; guard refusal without env config; serialization-latency capture alongside counts (decision 5.8 reports tokens AND latency). - Spending API credits is out of scope entirely — the maintainer runs the live phase; your job ends at a correct, guarded, fake-counter-tested integration.
- Validation gate:
uv run ruff check benchmarks/,uv run pyright benchmarks/,uv run pytest tests/benchmarks -q, fulluv run pytest --tb=short -q,uv run python-docs-mcp-server doctor.
4. Known pitfalls
- D6b sanction scope: registering
"anthropic"inPROVIDER_API_KEY_ENVinvalidates the merged test asserting"anthropic"is an unknown provider (tests/benchmarks/test_adapters.py:141-147). The sanction covers EXACTLY re-pointing that one assertion at a different non-anthropic unknown-provider string — nothing else. Any other merged-test change = stop-and-comment. - Dependency need = stop-and-comment (pipeline §8): any dependency addition is a maintainer-owned commit followed by re-pre-flight. Never commit it yourself.
- PR body must include
git diff --stat main...HEADprovingpyproject.tomlanduv.lockare absent from the diff. - Runtime network is a supervisor trigger: the server's offline-first runtime (principle 2.2) is untouchable — network lives only in the guarded benchmark live path, and tests must prove structural incapability elsewhere.
Refs #63, neverCloses #63.- Soft-ordered after [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86 merges (OPEN at draft time) — shared files:
benchmarks/adapters/__init__.py,guard.py,report.py. Rebase if [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86 lands mid-work. - Branch:
agent/89-claude-token-integration.
5. Decision log (worker fills in)
- …
- addedagent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agent
on Jul 8, 2026
Context
Parent: #63. Methodology:
docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md(work package 6; "Token Measurement" section). Roadmap decision 5.8: Claude tokenizer, measured after client-side rewrap. Maintainer decision on record (PLAN.md Amendment 2026-07-08): the Anthropic count-tokens API is the counting mechanism, confined to the maintainer-run live phase — exact counts for headline claims, zero runtime network access or new dependencies in the server itself. Token records (merged #75) carryclient_wrapped_tokens: None/raw_payload_tokens: Noneplaceholders;benchmarks/model_matrix.pyexposesMETHODOLOGY_TOKEN_LABEL = "Claude Tokens (Normalized Payload)".Status: NOT agent-ready. Filed by the orchestrator per PLAN.md T6(b); Vision must review, create the
.planning/agent-context/file (decision 5.14, mirrored as an issue comment), and applyagent-readybefore any dispatch.Goal
Add the Claude token-count integration: capture the client-rewrapped message envelope per cell, count it via the Anthropic count-tokens API behind the live-phase guard, and fill the runner's token records — with approximations clearly marked and barred from headline claims.
Acceptance criteria (draft — Vision edits before agent-ready)
approximation: trueand the report generator ([v0.5.0] benchmark reporting — generate raw report and README-safe summary #74) must exclude it from headline-eligible tables.require_live_environment()(same guard as [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #85) plus an Anthropic key env var; CI and unit tests never call it — tests use a fake counter.client_wrapped_tokensandraw_payload_tokens(raw payload counted separately as diagnostic data), labeledClaude Tokens (Normalized Payload)via the exported constant — never confusable with provider billing tokens.anthropic), if proposed, is a supervisor-review trigger: justify why a direct HTTPS call via existing deps is insufficient, and in any case the dependency must be an optional extra or dev/benchmark-only group — never a server runtime dependency.tests/benchmarks/) cover fake-counter records, approximation marking, guard refusal without env config, and serialization-latency capture alongside counts (decision 5.8 reports tokens AND latency).Refs #63, neverCloses #63.Scope boundaries
In scope: counting integration, guard, record filling, fake-counter tests. Out of scope: spending API credits (maintainer live phase only), report layout (#74), corpus, README.
Forbidden-territory reminder
Standard §2 list applies (notably:
pyproject.toml[project]changes and new runtime deps are supervisor-review at minimum); merged tests additive-only; the server's offline-first runtime (principle 2.2) is untouchable — network lives only in the guarded benchmark live path.Validation commands
uv run ruff check src/ tests/ uv run ruff check benchmarks/ uv run pyright src/ uv run pyright benchmarks/ uv run pytest --tb=short -q uv run pytest tests/benchmarks -q uv run python-docs-mcp-server doctor uv lock --check # if any dependency/extras change is proposedRecovery
If client-rewrap capture proves impossible for a given client, stop and comment here with the exact gap — the methodology requires marking that result as an approximation, not guessing.
Effort estimate
4–8 hours.
Refs #63.