Part of #16228, discovered while implementing #16230.
What's missing
services/voice_realtime_telemetry.py::_estimate_cost prices a Realtime WebRTC session from the live pricing store (llm_shared.pricing.redis_store.PricingRedisStore, #16229), which carries per-TOKEN rates (input_per_1m/output_per_1m/cache_read_per_1m). OpenAI's Realtime API also bills audio-seconds at a separate per-minute rate (previously a hardcoded _AUDIO_COST_PER_SEC dict, removed by #16230). Neither LiteLLM's nor OpenRouter's catalogue publishes a per-second audio rate, so there is currently no live source for it.
_estimate_cost therefore only prices the token component; the audio-seconds component is omitted rather than estimated from a hardcoded rate. This means:
RealtimeSessionRecord.estimated_cost_usd (and the cost-cap check in check_caps) undercount a session's real spend for any session with meaningful audio-only usage (e.g. a long voice call with few tokens).
- The cost cap (
AUTOBOT_VOICE_REALTIME_MAX_COST_USD) can no longer be trusted to catch an audio-heavy runaway session -- it now only bounds token spend.
Suggested scope
Where
autobot-backend/services/voice_realtime_telemetry.py::_estimate_cost and check_caps.
Part of #16228, discovered while implementing #16230.
What's missing
services/voice_realtime_telemetry.py::_estimate_costprices a Realtime WebRTC session from the live pricing store (llm_shared.pricing.redis_store.PricingRedisStore, #16229), which carries per-TOKEN rates (input_per_1m/output_per_1m/cache_read_per_1m). OpenAI's Realtime API also bills audio-seconds at a separate per-minute rate (previously a hardcoded_AUDIO_COST_PER_SECdict, removed by #16230). Neither LiteLLM's nor OpenRouter's catalogue publishes a per-second audio rate, so there is currently no live source for it._estimate_costtherefore only prices the token component; the audio-seconds component is omitted rather than estimated from a hardcoded rate. This means:RealtimeSessionRecord.estimated_cost_usd(and the cost-cap check incheck_caps) undercount a session's real spend for any session with meaningful audio-only usage (e.g. a long voice call with few tokens).AUTOBOT_VOICE_REALTIME_MAX_COST_USD) can no longer be trusted to catch an audio-heavy runaway session -- it now only bounds token spend.Suggested scope
model_prices_and_context_window.jsonor OpenRouter's/api/v1/modelsas far as pricing(sources): fetch live prices from LiteLLM and OpenRouter, cross-checked, with honest freshness #16229's sources go, so this may need a third catalogue, or an operator-set override via the existingPUT /api/admin/pricing/{provider}/{model}mechanism with a dedicated audio-rate field.ModelPricing(or a sibling model) with an audio-rate dimension once a source exists.check_capsshould treat a Realtime session with nonzero audio time as "cost unknown" too (matching pricing(remove): delete every hardcoded price and price table, rewire consumers, unknown is never $0 #16230's "never $0, never guessed" rule) rather than silently reporting a token-only figure as the session's cost.Where
autobot-backend/services/voice_realtime_telemetry.py::_estimate_costandcheck_caps.