Skip to content

bug(optimization): the KV sliding window has no caller, and its tail-only policy would evict the prompt prefix (#13030) #17801

Description

@mrveiss

Context — #13033's defect is fixed; this is what is left

#13030 D4 said trim_to_length() retained the oldest window instead of the newest. That is
fixed in merged code:

$ git show origin/main:autobot-backend/llm_shared/optimization/kv_cache.py | sed -n '291,305p'
    @staticmethod
    def _retain_tail(entry: "_LayerEntry", max_len: int) -> None:
        """Move the newest ``max_len`` positions of *entry* down to offset 0.
        ...
        start = entry.filled_len - max_len
        if max_len > 0:
            entry.k[:, :max_len, :, :].copy_(entry.k[:, start : entry.filled_len, :, :].clone())
            entry.v[:, :max_len, :, :].copy_(entry.v[:, start : entry.filled_len, :, :].clone())
        entry.filled_len = max_len

Two separate problems remain, and neither is in #13033's text.

Problem 1 — the sliding window has no caller

$ grep -rn "trim_to_length" --include=*.py . | grep -v __pycache__
autobot-backend/llm_shared/optimization/kv_cache_test.py:  (13 hits)
autobot-backend/llm_shared/optimization/kv_cache.py:254:    def trim_to_length(...)

Zero production callers. The window has been specified, implemented, broken, corrected and tested,
and has never run. Per the no-deletion rule this is unfinished work to wire in, not dead code.

Problem 2 — a tail-only policy is the wrong policy for the thing that will call it

_retain_tail keeps the last max_len positions, unconditionally. The first caller will be a
decode loop whose sequence begins with a system prompt, a task description and (on any multimodal
path) an encoded input prefix. Wiring the window as it stands evicts exactly those — the tokens
that must survive — and does so silently, which is the same failure shape #13033 was filed for.

The established pattern for this is to split retention in two: a pinned prefix that is written
once and never evicted, and a fixed ring over the positions after it, where each new position
overwrites prefix_len + (pos % W) in place and a modulo counter advances. Memory is then constant
in output length, and the prefix is structurally safe rather than safe by convention.

It is also cheaper. _retain_tail clone()s the retained window and copy_s it down to offset 0
on every trim — O(window) per call. A ring writes one slot — O(1).

Note the cost this buys, so it is adopted with eyes open: ring order is not chronological, so
get()'s k[:, :filled_len] contract, any mask construction and any cache-inspection code must be
re-checked against it. A windowed cache also cannot see what it emitted more than W positions
ago, which is a real generation-quality trade, not a free win. A token-plane n-gram guard is the
usual compensation and is deliberately out of scope here — #13892 evaluated and deferred it on
the grounds that nothing in AutoBot generates long-form output over documents today.

Acceptance criteria

  • LayerKVCache supports a pinned-prefix + ring retention mode alongside _retain_tail, with
    the prefix length set once per sequence and never evicted.
  • Ring writes are in place (one slot per position), with no per-step clone or copy-down.
  • get()'s ordering contract is stated explicitly for both modes, and every consumer of it is
    checked against the ring mode.
  • trim_to_length (or the ring mode) has a production caller, or this issue states in writing
    why it still does not and what blocks it.
  • A test asserts the pinned prefix survives an output long enough to wrap the ring more than
    once.
  • A test asserts KV memory is constant across output lengths that differ by an order of
    magnitude.

Blocked by

#13031 — forward_pass accepts a kv_cache argument and documents that it does not inspect it
(layer_inference.py:406-408),
and generate() keeps no cache at all. A retention policy is unverifiable until something caches.

Provenance

docs/research/long-horizon-document-parsing-ring-cache.md,
"What We Can Adopt" item 1 and "Gaps & Opportunities" item 3.

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions