Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .changeset/fable-cache-read-multiplier.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"@inbrace-tech/tokenline": patch
---

fix: bill cache reads at 0.025x on Claude Fable 5.1 / Mythos 5.1. The per-turn economics line used a flat 0.1x cache-read multiplier for every Claude model, overstating `eq` and understating `saving %` on the models with the discounted cache-hit price. `tokenline.sh` now reads `model.id` and picks 0.025x for Fable 5.1 / Mythos 5.1; every other model keeps 0.1x.
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,7 +67,7 @@ LLMs are stateless — by default, they must reread your entire codebase context

**Prompt Caching** solves this by saving that processed state in the provider's memory:
- **Cache Write (Cold):** The model reads everything and stores the state. This costs slightly more than base tokens.
- **Cache Hit (Warm):** If your prompt prefix matches the cached state exactly, the model skips the reading phase. This is **~90% cheaper** and starts responding almost instantly.
- **Cache Hit (Warm):** If your prompt prefix matches the cached state exactly, the model skips the reading phase. This is **~90% cheaper** (**97.5%** on Claude Fable 5.1 / Mythos 5.1, which bill cache hits at 0.025x base input) and starts responding almost instantly. `tokenline` picks the right multiplier from the active model.
- **TTL (Time-To-Live):** The cache operates as a **sliding window**. Every cache hit resets the countdown to its full duration for free. The secret to minimizing rate-limit drain is adjusting your workflow pace to keep the cache continuously **HOT**.

### 💡 Pro-tip: Forcing a 5m TTL to save rate limits
Expand Down
14 changes: 13 additions & 1 deletion tokenline.sh
Original file line number Diff line number Diff line change
Expand Up @@ -95,7 +95,8 @@ parse_and_prepare_paths() {
(.context_window.current_usage.input_tokens // 0),
(.context_window.current_usage.output_tokens // 0),
(.context_window.current_usage.cache_creation_input_tokens // 0),
(.context_window.current_usage.cache_read_input_tokens // 0)' 2>/dev/null)
(.context_window.current_usage.cache_read_input_tokens // 0),
(.model.id // "")' 2>/dev/null)

# Malformed or empty stdin: jq emits nothing, so the array is empty. Degrade to
# a silent no-op render rather than leaking parse errors or rendering garbage —
Expand All @@ -115,6 +116,7 @@ parse_and_prepare_paths() {
cur_output="${_f[10]}"
cur_cwrite="${_f[11]}"
cur_cread="${_f[12]}"
model_id="${_f[13]}"

# Computed: total input-only tokens used in the current context window
tokens_used=$((cur_input + cur_cwrite + cur_cread))
Expand All @@ -141,6 +143,15 @@ parse_and_prepare_paths() {
is_gemini=true
fi

# Claude Fable 5.1 and Mythos 5.1 bill cache hits at 0.025x base input instead of
# the standard 0.1x. Match the API id (claude-fable-5-1[1m]) or the display name
# (Fable 5.1) so either field alone is enough.
discounted_cache_read=false
local discounted_re='(fable|mythos)[- ]5[-.]1([^0-9]|$)'
if [[ "${model_id,,}" =~ $discounted_re ]] || [[ "${model,,}" =~ $discounted_re ]]; then
discounted_cache_read=true
fi

# Get the current epoch timestamp once to be reused across all calculations
now=$(date +%s)
}
Expand Down Expand Up @@ -402,6 +413,7 @@ compute_turn_breakdown() {
output_mult="4"
else
read_mult="0.1"
[ "$discounted_cache_read" = true ] && read_mult="0.025"
write_mult="1.25"
[ "${ttl_label:-5m}" = "1h" ] && write_mult="2"
input_mult="1"
Expand Down