Skip to content

Fix/xdg isolation kv cache flush permissions - #4

Merged
cebeach merged 4 commits into
mainfrom
fix/xdg-isolation-kv-cache-flush-permissions-v2
Mar 14, 2026
Merged

cebeach merged 4 commits into
mainfrom
fix/xdg-isolation-kv-cache-flush-permissions-v2

Conversation

@cebeach

@cebeach cebeach commented Mar 14, 2026

Copy link
Copy Markdown
Owner

feat(infra): llama-server VRAM analysis, GGUF metadata inspection, and optimized server configuration

Overview

Three new tools and an updated server launch script, developed through
iterative empirical analysis of gpt-oss-20b MXFP4 running under
llama-server on an RTX 4090. The work uncovered non-obvious behavior
around KV cache allocation, slot concurrency, and context length
management that informed the final server configuration.

llama_server_vram_report.py

A log file parser that extracts and pretty-prints a structured VRAM
metrics report from a llama-server startup log.

Sections reported:

  • Model: name, architecture, file type, size, quantization, context
    length, KV quantization
  • Model architecture: layer count, full-attention vs SWA layer split,
    SWA window size, embedding dimension, attention heads (Q and KV),
    head dimensions, MoE expert count and active experts per token,
    expert utilization percentage, RoPE scaling type and factor
  • VRAM allocation: total VRAM, free at startup, per-component
    breakdown (model weights CUDA, CPU mapped, KV cache non-SWA, KV
    cache SWA, compute buffer), total used, free after load
  • KV cache detail: slot count, kv_unified flag, context per slot
    (with capping warning when --ctx-size exceeds model training
    context), non-SWA and SWA cell counts and layer counts, per-slot
    KV cost
  • Slot capacity analysis: free VRAM after load, llama.cpp safety
    margin, usable headroom, additional slots possible

Key findings documented through empirical log analysis:

KV cache is O(ctx_size), compute buffer is also O(ctx_size).
The compute buffer grows approximately linearly with --ctx-size at
~768 MiB per 131,072 tokens, making it the binding constraint when
scaling context — not the KV cache itself. The OOM boundary for
gpt-oss-20b MXFP4 on RTX 4090 lies between --ctx-size 524,288 and
655,360 at --parallel 16 --kv-unified.

kv_unified = true is not the default when --parallel is set
explicitly.
When --parallel is omitted, the server auto-selects
n_parallel = 4 and sets kv_unified = true (hardcoded in server.cpp).
When --parallel N is passed explicitly, kv_unified defaults to false,
causing the pool to be divided equally across slots (ctx_size /
n_parallel tokens per slot). The --kv-unified flag must be passed
explicitly alongside --parallel to preserve full per-slot context.

In unified mode, slot count is essentially free. The KV pool size
is fixed by --ctx-size regardless of --parallel. Adding slots from 4
to 8 to 16 produced identical total KV allocation (1,632 MiB) with
negligible overhead. The per-slot cost figure in the capacity analysis
is therefore misleading in unified mode and should be interpreted as
a scheduling budget, not a memory budget.

--ctx-size is capped per slot at the model's training context.
Setting --ctx-size 524,288 with --parallel 4 does not give 4 slots
at 512K context. llama-server caps each slot at the model's training
context length (131,072 for gpt-oss-20b), logging a warning:
"the slot context (524288) exceeds the training context of the model
(131072) - capping". The pool holds 524,288 cells total but each slot
is limited to 131,072. The report surfaces this via a "Context per
slot" row with a capping note.

Bug fixes during development:

  • Slot count regex now tries initializing slots, n_slots = N first
    (present when --parallel is explicit), then the auto-detection
    format, then the bare n_parallel fallback. Previously only the
    auto-detection format was checked, causing explicit --parallel
    invocations to report 1 slot.
  • SWA/non-SWA layer counts deduplicate across the two log passes
    (fit probe + actual load) to avoid double-counting.
  • CUDA model buffer uses the last non-zero occurrence to skip the
    fit probe's 0.00 MiB entry.

gguf_model_info.py

A standalone GGUF metadata inspector using the gguf and prettytable
packages. Reads architecture metadata directly from the model file
without requiring a running server.

Sections reported in Unicode box-drawing tables:

  • General: name, architecture, size label, file type, license
  • Context & dimensions: max context length, layer count, embedding
    dimension, feed-forward length, vocabulary size
  • Attention: Q heads, KV heads, GQA ratio, head dimensions, SWA
    window size and pattern, RMS norm epsilon
  • Mixture of experts: total experts, active experts per token,
    utilization percentage, expert groups, shared experts, expert
    FF length
  • RoPE positional encoding: frequency base, dimension count,
    scaling type and factor, original context length, context
    extension multiplier
  • Tokenizer: model, pre, BOS/EOS/PAD token IDs
  • Tensors: total tensor count, total parameter count, per-
    quantization-type breakdown with tensor count and parameter
    count with percentage of total

The tensor breakdown for gpt-oss-20b MXFP4 confirms the naming:
MXFP4 accounts for 91.4% of parameters (72 tensors, expert weight
matrices), Q8_0 accounts for 8.6% (98 tensors, attention weights),
and F32 accounts for 0.04% (289 tensors, norm weights and embeddings).
Despite the majority of tensors being F32, they represent a negligible
fraction of total parameters.

Uses Keys constants from the gguf package for field name resolution,
making the script architecture-agnostic — the {arch} template is
resolved at runtime from general.architecture.

Requires: pip install gguf prettytable

llama-server-openai_gpt-oss-20b-MXFP4.sh (updated)

Updated server launch script incorporating findings from the VRAM
analysis work.

Changes from prior version:

  • Added --kv-unified to ensure the shared context pool is maintained
    when --parallel is specified explicitly
  • Added --parallel ${SLOTS} with SLOTS=4 as a named variable
  • Added --ctx-size ${CTX_SIZE} computed as SLOTS * CTX_LEN, making
    the relationship between slot count and pool size explicit and
    self-documenting
  • Added --cache-type-k q8_0 --cache-type-v q8_0 for KV quantization,
    halving KV cache VRAM vs f16 (408 MiB/slot vs 768 MiB/slot)
  • Added --slot-save-path ${TMP}/slots, required to enable the
    POST /slots/{id}?action=erase endpoint used by the pytest harness
    KV cache flush fixture
  • Moved log files and slot state to ${TMP}/log and ${TMP}/slots
    under a shared TMP=/tmp/llama-server root
  • Added llama-server --version capture to log header using
    $(llama-server --version 2>&1 | tr -s ' \n' ' ' | xargs)
  • Added commentary on reasoning level configuration (low/medium/high
    set via system prompt "Reasoning: ") and the slot/KV cache
    architecture informed by source analysis
  • CTX_LEN and SLOTS promoted to named variables at the top of the
    script for easy reconfiguration

Resultant configuration on RTX 4090 with gpt-oss-20b MXFP4:

  • 4 slots, each with full 131,072-token context (unified pool)
  • Total KV cache: ~1,657 MiB (1,632 non-SWA + 25 SWA at Q8_0)
  • Total VRAM used: ~13,448 MiB
  • Free after load: ~8,500 MiB
  • llama.cpp safety margin: 1,024 MiB
  • Effective headroom: ~7,476 MiB

llama_server_slots_brief_r1.md

A technical reference document covering the slot system architecture,
KV cache allocation modes, the concurrency model and its distinction
from CPU thread parallelism, the KV cache flush API, and the
relationship to OpenCode's parallel agent roadmap. Intended as a
reference for contributors working with llama-server in multi-session
or multi-agent contexts.

cebeach added 4 commits March 13, 2026 20:11
…grant, and regex assertions

Resolve four categories of test reliability issues identified through
iterative debugging of the black-box pytest harness against gpt-oss-20b
running under llama-server.

XDG directory isolation
- Move .xdg_config, .xdg_state, .xdg_cache from inside workspace/ to
  siblings under tmp_path, preventing OpenCode's bun cache and
  node_modules from polluting snapshot_text_files() and causing
  spurious files_unchanged assertion failures
- Add XDG_DATA_HOME redirection to .xdg_data/ — without this, session
  files written to ~/.local/share/opencode leaked between runs, causing
  the second run to find a stale session referencing the previous
  tmp_path and triggering a permission dialog against a non-existent
  workspace (anomalyco/opencode#8538 root cause)
- Document the tmp_path layout in a block comment above xdg_base

KV cache flush fixture
- Add flush_llama_kv_cache autouse fixture to conftest.py, erasing all
  llama-server slots via POST /slots/{id}?action=erase&id_slot={id}
  before each test
- Without this, residual KV cache from prior tests caused the model to
  reference stale workspace paths from earlier tmp_path directories
- Slot IDs are discovered dynamically from GET /slots so the fixture
  adapts to any --parallel value
- Failure is a warning not an error — flush is best-effort
- Requires llama-server to be started with --slot-save-path (gates the
  POST /slots endpoint regardless of whether save/restore is used)
- Add requests to requirements.txt

Permission pre-grant
- Add permission block to tools/opencode/opencode.json granting allow
  for read, glob, grep, edit, write, and external_directory
- Eliminates the interactive permission dialog on every invocation:
  OpenCode raises external_directory for all workspace paths because the
  ephemeral workspace has no .git root, causing project ID to resolve
  to "global" (anomalyco/opencode#8538)
- Permission state is in-memory only and not persisted across processes
  (confirmed in permission/next.ts), so dialog would fire on every run
  regardless of prior approvals
- Permission approval fallback (Tab+Enter for Allow always) retained in
  run_opencode_tui with saw_post_submit_output reset after each approval
  to prevent READY_RE match on post-approval redraw from prematurely
  terminating the response-wait loop

Regex output assertions
- Add assert_matches_all and assert_not_matches_any functions accepting
  regex patterns, wired to new YAML keys stdout_contains_re,
  stdout_not_contains_re, stderr_contains_re, stderr_not_contains_re
- Required because TUI terminal wrapping splits long paths across lines,
  breaking literal stdout_contains assertions for nested paths
- Update glob_discovery.yaml to use stdout_contains_re with pattern
  src[\s/]*nested[\s/]*sample\.py, matching the path regardless of
  how the TUI wraps it across line boundaries

Other
- Increase quiet_timeout from 2s to 5s to allow TUI rendering to
  complete before the harness declares the response done
- Update write_file.yaml prompt to explicitly invoke the write tool,
  preventing the model from responding with a plan instead of executing
- Add requests and pexpect to README install instructions
@cebeach
cebeach merged commit 55b556b into main Mar 14, 2026
@cebeach
cebeach deleted the fix/xdg-isolation-kv-cache-flush-permissions-v2 branch March 14, 2026 07:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant