Skip to content

[v0.5.0] benchmark tokens — Claude token-count integration after client rewrap (live-phase-gated) #89

Description

@ayhammouda

Context

Parent: #63. Methodology: docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md (work package 6; "Token Measurement" section). Roadmap decision 5.8: Claude tokenizer, measured after client-side rewrap. Maintainer decision on record (PLAN.md Amendment 2026-07-08): the Anthropic count-tokens API is the counting mechanism, confined to the maintainer-run live phase — exact counts for headline claims, zero runtime network access or new dependencies in the server itself. Token records (merged #75) carry client_wrapped_tokens: None / raw_payload_tokens: None placeholders; benchmarks/model_matrix.py exposes METHODOLOGY_TOKEN_LABEL = "Claude Tokens (Normalized Payload)".

Status: NOT agent-ready. Filed by the orchestrator per PLAN.md T6(b); Vision must review, create the .planning/agent-context/ file (decision 5.14, mirrored as an issue comment), and apply agent-ready before any dispatch.

Goal

Add the Claude token-count integration: capture the client-rewrapped message envelope per cell, count it via the Anthropic count-tokens API behind the live-phase guard, and fill the runner's token records — with approximations clearly marked and barred from headline claims.

Acceptance criteria (draft — Vision edits before agent-ready)

  • Envelope capture: the counting path takes the same client-side wrapping used by the benchmark client (methodology "Token Measurement" steps 1–3); where a client cannot expose its exact wrapped envelope, the record is marked approximation: true and the report generator ([v0.5.0] benchmark reporting — generate raw report and README-safe summary #74) must exclude it from headline-eligible tables.
  • The count-tokens API call sits behind require_live_environment() (same guard as [v0.5.0] benchmark adapters — define OpenAI/Google model matrix #85) plus an Anthropic key env var; CI and unit tests never call it — tests use a fake counter.
  • Completed token records fill client_wrapped_tokens and raw_payload_tokens (raw payload counted separately as diagnostic data), labeled Claude Tokens (Normalized Payload) via the exported constant — never confusable with provider billing tokens.
  • A new third-party runtime dependency (e.g. anthropic), if proposed, is a supervisor-review trigger: justify why a direct HTTPS call via existing deps is insufficient, and in any case the dependency must be an optional extra or dev/benchmark-only group — never a server runtime dependency.
  • Tests (additive, under tests/benchmarks/) cover fake-counter records, approximation marking, guard refusal without env config, and serialization-latency capture alongside counts (decision 5.8 reports tokens AND latency).
  • No README/benchmark claims; Refs #63, never Closes #63.

Scope boundaries

In scope: counting integration, guard, record filling, fake-counter tests. Out of scope: spending API credits (maintainer live phase only), report layout (#74), corpus, README.

Forbidden-territory reminder

Standard §2 list applies (notably: pyproject.toml [project] changes and new runtime deps are supervisor-review at minimum); merged tests additive-only; the server's offline-first runtime (principle 2.2) is untouchable — network lives only in the guarded benchmark live path.

Validation commands

uv run ruff check src/ tests/
uv run ruff check benchmarks/
uv run pyright src/
uv run pyright benchmarks/
uv run pytest --tb=short -q
uv run pytest tests/benchmarks -q
uv run python-docs-mcp-server doctor
uv lock --check   # if any dependency/extras change is proposed

Recovery

If client-rewrap capture proves impossible for a given client, stop and comment here with the exact gap — the methodology requires marking that result as an approximation, not guessing.

Effort estimate

4–8 hours.

Refs #63.

Activity

  1. added this to the v0.5.0 milestone on Jul 8, 2026
  2. ayhammouda commented on Jul 8, 2026

    @ayhammouda
    OwnerAuthor

    Pre-flight amendments (2026-07-08, maintainer-approved — sign-off mirrored on #63)

    Two binding changes to this issue's draft criteria:

    1. Dependency rule TIGHTENED (supersedes the body's "optional extra" language; Codex round-1 finding 2): the worker may NOT add any dependency — no anthropic SDK, no optional extra, no dev/benchmark group. pyproject.toml and uv.lock must be absent from the diff entirely; the PR body must include git diff --stat main...HEAD proving it. Implement the count-tokens call via the standard library or already-installed dependencies, strictly behind the live-phase guard. If an SDK appears genuinely required: stop and comment (pipeline §8); any dependency addition is a maintainer-owned commit followed by re-pre-flight.

    2. Named merged-test sanction (D6b): registering "anthropic": "ANTHROPIC_API_KEY" in PROVIDER_API_KEY_ENV (benchmarks/adapters/guard.py:28-31) invalidates the merged test asserting "anthropic" is an unknown provider (tests/benchmarks/test_adapters.py). The sanction covers exactly: re-pointing that one assertion at a different unknown-provider string. Nothing else — any other merged-test change remains stop-and-comment.

    Refs #63

  3. ayhammouda commented on Jul 8, 2026

    @ayhammouda
    OwnerAuthor

    Agent context — #89 [v0.5.0] benchmark tokens — Claude token-count integration after client rewrap (live-phase-gated)

    Status: DRAFT for dispatch (verified against main @ ab9e615, 2026-07-08)
    Auditable copy: this content is mirrored as a comment on issue #89 (.planning/ is gitignored — the issue comment is the copy pre-flight verifies and the dispatch packet links).

    1. Roadmap excerpt (why this exists)

    • Locked decision (PLAN Amendment 1, restated in [v0.5.0] benchmark tokens — Claude token-count integration after client rewrap (live-phase-gated) #89): the Anthropic count-tokens API is the counting mechanism, confined to the maintainer-run live phase. Exact counts back headline claims; the server itself gains zero runtime network access and zero new dependencies.
    • Methodology "Token Measurement" steps 1–3 (docs/benchmarks/PUBLIC-BENCHMARK-METHODOLOGY.md): count the envelope after client-side rewrap; where a client cannot expose its exact wrapped envelope, mark the record approximation: true — approximations are never headline-eligible.
    • BINDING pre-flight amendment (issue [v0.5.0] benchmark tokens — Claude token-count integration after client rewrap (live-phase-gated) #89 comment, 2026-07-08 — supersedes the body's "optional extra" language): the worker may NOT add any dependency — no anthropic SDK, no optional extra, no dev/benchmark group. pyproject.toml and uv.lock must be absent from the diff entirely. Implement the count-tokens call via the standard library or already-installed dependencies, strictly behind the live-phase guard.
    • The same comment grants the D6b merged-test sanction — scope in section 4. Both amendments are binding on this dispatch.

    2. Code touch-points (verify against merged main before starting)

    • benchmarks/adapters/guard.py:54-81 — require_live_environment(provider): the count-tokens call MUST sit behind this guard (flag + key, both required).
    • benchmarks/adapters/guard.py:25 — LIVE_PROVIDERS_ENABLED_ENV = "BENCHMARK_LIVE_PROVIDERS_ENABLED": the explicit opt-in flag; reuse it, do not invent a second flag.
    • benchmarks/adapters/guard.py:28-31 — PROVIDER_API_KEY_ENV: register "anthropic": "ANTHROPIC_API_KEY" here.
    • benchmarks/adapters/guard.py:43-51 — LiveExecutionNotImplementedError: the existing pattern for "guard passed but structurally incapable"; follow it for any path CI could reach.
    • benchmarks/adapters/base.py:49-64 — TokenMetadata: provider-native counts are diagnostic only and must never be reported under METHODOLOGY_TOKEN_LABEL.
    • benchmarks/model_matrix.py:33 — METHODOLOGY_TOKEN_LABEL = "Claude Tokens (Normalized Payload)": the only label for methodology counts; use the exported constant, never a string literal.
    • benchmarks/runner.py:250-260 — token_record dict; client_wrapped_tokens / raw_payload_tokens placeholders are None at :257-258. Fill both (raw payload counted separately, diagnostic). Update the stale notes string at :259 ("out of scope for issue [v0.5.0] benchmark runner — add reproducible CLI and artifact layout #72").
    • benchmarks/report.py — CellRecord token fields :102-103, parsed :223-224; _basic_stats token aggregation :355-379; family headline_eligible rendering :480-481. Additive change: exclude approximation: true records from headline-eligible tables.
    • Count-tokens HTTP call: stdlib (urllib.request) or already-installed deps ONLY, invoked strictly after require_live_environment("anthropic") passes.

    3. Test patterns to follow

    • Guard env monkeypatch pattern: tests/benchmarks/test_adapters.py:106-148 — monkeypatch.delenv/setenv on LIVE_PROVIDERS_ENABLED_ENV + PROVIDER_API_KEY_ENV[...], assert LiveProviderDisabledError with match=.
    • Use a fake counter for ALL tests; CI and unit tests never call the real API. Prove structurally that no test path reaches the network.
    • Cover, additively under tests/benchmarks/: fake-counter record filling; approximation marking + headline exclusion; guard refusal without env config; serialization-latency capture alongside counts (decision 5.8 reports tokens AND latency).
    • Spending API credits is out of scope entirely — the maintainer runs the live phase; your job ends at a correct, guarded, fake-counter-tested integration.
    • Validation gate: uv run ruff check benchmarks/, uv run pyright benchmarks/, uv run pytest tests/benchmarks -q, full uv run pytest --tb=short -q, uv run python-docs-mcp-server doctor.

    4. Known pitfalls

    • D6b sanction scope: registering "anthropic" in PROVIDER_API_KEY_ENV invalidates the merged test asserting "anthropic" is an unknown provider (tests/benchmarks/test_adapters.py:141-147). The sanction covers EXACTLY re-pointing that one assertion at a different non-anthropic unknown-provider string — nothing else. Any other merged-test change = stop-and-comment.
    • Dependency need = stop-and-comment (pipeline §8): any dependency addition is a maintainer-owned commit followed by re-pre-flight. Never commit it yourself.
    • PR body must include git diff --stat main...HEAD proving pyproject.toml and uv.lock are absent from the diff.
    • Runtime network is a supervisor trigger: the server's offline-first runtime (principle 2.2) is untouchable — network lives only in the guarded benchmark live path, and tests must prove structural incapability elsewhere.
    • Refs #63, never Closes #63.
    • Soft-ordered after [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86 merges (OPEN at draft time) — shared files: benchmarks/adapters/__init__.py, guard.py, report.py. Rebase if [v0.5.0] benchmark adapter — python-docs-mcp-server tool runner (offline stdio) #86 lands mid-work.
    • Branch: agent/89-claude-token-integration.

    5. Decision log (worker fills in)

    • …
  4. added
    agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agent
    on Jul 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agent-readyIssue passed AGENT-EXECUTION-PIPELINE.md §10 pre-flight; scoped for an autonomous agentenhancementNew feature or requestpriority:P2Medium priority

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions