Context — #13033's defect is fixed; this is what is left
#13030 D4 said trim_to_length() retained the oldest window instead of the newest. That is
fixed in merged code:
$ git show origin/main:autobot-backend/llm_shared/optimization/kv_cache.py | sed -n '291,305p'
@staticmethod
def _retain_tail(entry: "_LayerEntry", max_len: int) -> None:
"""Move the newest ``max_len`` positions of *entry* down to offset 0.
...
start = entry.filled_len - max_len
if max_len > 0:
entry.k[:, :max_len, :, :].copy_(entry.k[:, start : entry.filled_len, :, :].clone())
entry.v[:, :max_len, :, :].copy_(entry.v[:, start : entry.filled_len, :, :].clone())
entry.filled_len = max_len
Two separate problems remain, and neither is in #13033's text.
Problem 1 — the sliding window has no caller
$ grep -rn "trim_to_length" --include=*.py . | grep -v __pycache__
autobot-backend/llm_shared/optimization/kv_cache_test.py: (13 hits)
autobot-backend/llm_shared/optimization/kv_cache.py:254: def trim_to_length(...)
Zero production callers. The window has been specified, implemented, broken, corrected and tested,
and has never run. Per the no-deletion rule this is unfinished work to wire in, not dead code.
Problem 2 — a tail-only policy is the wrong policy for the thing that will call it
_retain_tail keeps the last max_len positions, unconditionally. The first caller will be a
decode loop whose sequence begins with a system prompt, a task description and (on any multimodal
path) an encoded input prefix. Wiring the window as it stands evicts exactly those — the tokens
that must survive — and does so silently, which is the same failure shape #13033 was filed for.
The established pattern for this is to split retention in two: a pinned prefix that is written
once and never evicted, and a fixed ring over the positions after it, where each new position
overwrites prefix_len + (pos % W) in place and a modulo counter advances. Memory is then constant
in output length, and the prefix is structurally safe rather than safe by convention.
It is also cheaper. _retain_tail clone()s the retained window and copy_s it down to offset 0
on every trim — O(window) per call. A ring writes one slot — O(1).
Note the cost this buys, so it is adopted with eyes open: ring order is not chronological, so
get()'s k[:, :filled_len] contract, any mask construction and any cache-inspection code must be
re-checked against it. A windowed cache also cannot see what it emitted more than W positions
ago, which is a real generation-quality trade, not a free win. A token-plane n-gram guard is the
usual compensation and is deliberately out of scope here — #13892 evaluated and deferred it on
the grounds that nothing in AutoBot generates long-form output over documents today.
Acceptance criteria
Blocked by
#13031 — forward_pass accepts a kv_cache argument and documents that it does not inspect it
(layer_inference.py:406-408),
and generate() keeps no cache at all. A retention policy is unverifiable until something caches.
Provenance
docs/research/long-horizon-document-parsing-ring-cache.md,
"What We Can Adopt" item 1 and "Gaps & Opportunities" item 3.
Context — #13033's defect is fixed; this is what is left
#13030 D4 said
trim_to_length()retained the oldest window instead of the newest. That isfixed in merged code:
Two separate problems remain, and neither is in #13033's text.
Problem 1 — the sliding window has no caller
Zero production callers. The window has been specified, implemented, broken, corrected and tested,
and has never run. Per the no-deletion rule this is unfinished work to wire in, not dead code.
Problem 2 — a tail-only policy is the wrong policy for the thing that will call it
_retain_tailkeeps the lastmax_lenpositions, unconditionally. The first caller will be adecode loop whose sequence begins with a system prompt, a task description and (on any multimodal
path) an encoded input prefix. Wiring the window as it stands evicts exactly those — the tokens
that must survive — and does so silently, which is the same failure shape #13033 was filed for.
The established pattern for this is to split retention in two: a pinned prefix that is written
once and never evicted, and a fixed ring over the positions after it, where each new position
overwrites
prefix_len + (pos % W)in place and a modulo counter advances. Memory is then constantin output length, and the prefix is structurally safe rather than safe by convention.
It is also cheaper.
_retain_tailclone()s the retained window andcopy_s it down to offset 0on every trim — O(window) per call. A ring writes one slot — O(1).
Note the cost this buys, so it is adopted with eyes open: ring order is not chronological, so
get()'sk[:, :filled_len]contract, any mask construction and any cache-inspection code must bere-checked against it. A windowed cache also cannot see what it emitted more than
Wpositionsago, which is a real generation-quality trade, not a free win. A token-plane n-gram guard is the
usual compensation and is deliberately out of scope here — #13892 evaluated and deferred it on
the grounds that nothing in AutoBot generates long-form output over documents today.
Acceptance criteria
LayerKVCachesupports a pinned-prefix + ring retention mode alongside_retain_tail, withthe prefix length set once per sequence and never evicted.
get()'s ordering contract is stated explicitly for both modes, and every consumer of it ischecked against the ring mode.
trim_to_length(or the ring mode) has a production caller, or this issue states in writingwhy it still does not and what blocks it.
once.
magnitude.
Blocked by
#13031 —
forward_passaccepts akv_cacheargument and documents that it does not inspect it(
layer_inference.py:406-408),and
generate()keeps no cache at all. A retention policy is unverifiable until something caches.Provenance
docs/research/long-horizon-document-parsing-ring-cache.md,"What We Can Adopt" item 1 and "Gaps & Opportunities" item 3.