Skip to content

pricing(voice): Realtime audio-seconds cost has no live source, so the cost cap only bounds token spend #16326

Description

@mrveiss

Part of #16228, discovered while implementing #16230.

What's missing

services/voice_realtime_telemetry.py::_estimate_cost prices a Realtime WebRTC session from the live pricing store (llm_shared.pricing.redis_store.PricingRedisStore, #16229), which carries per-TOKEN rates (input_per_1m/output_per_1m/cache_read_per_1m). OpenAI's Realtime API also bills audio-seconds at a separate per-minute rate (previously a hardcoded _AUDIO_COST_PER_SEC dict, removed by #16230). Neither LiteLLM's nor OpenRouter's catalogue publishes a per-second audio rate, so there is currently no live source for it.

_estimate_cost therefore only prices the token component; the audio-seconds component is omitted rather than estimated from a hardcoded rate. This means:

  • RealtimeSessionRecord.estimated_cost_usd (and the cost-cap check in check_caps) undercount a session's real spend for any session with meaningful audio-only usage (e.g. a long voice call with few tokens).
  • The cost cap (AUTOBOT_VOICE_REALTIME_MAX_COST_USD) can no longer be trusted to catch an audio-heavy runaway session -- it now only bounds token spend.

Suggested scope

Where

autobot-backend/services/voice_realtime_telemetry.py::_estimate_cost and check_caps.

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions