Repository navigation
Fix/xdg isolation kv cache flush permissions - #4
Merged
Merged
Conversation
…grant, and regex assertions Resolve four categories of test reliability issues identified through iterative debugging of the black-box pytest harness against gpt-oss-20b running under llama-server. XDG directory isolation - Move .xdg_config, .xdg_state, .xdg_cache from inside workspace/ to siblings under tmp_path, preventing OpenCode's bun cache and node_modules from polluting snapshot_text_files() and causing spurious files_unchanged assertion failures - Add XDG_DATA_HOME redirection to .xdg_data/ — without this, session files written to ~/.local/share/opencode leaked between runs, causing the second run to find a stale session referencing the previous tmp_path and triggering a permission dialog against a non-existent workspace (anomalyco/opencode#8538 root cause) - Document the tmp_path layout in a block comment above xdg_base KV cache flush fixture - Add flush_llama_kv_cache autouse fixture to conftest.py, erasing all llama-server slots via POST /slots/{id}?action=erase&id_slot={id} before each test - Without this, residual KV cache from prior tests caused the model to reference stale workspace paths from earlier tmp_path directories - Slot IDs are discovered dynamically from GET /slots so the fixture adapts to any --parallel value - Failure is a warning not an error — flush is best-effort - Requires llama-server to be started with --slot-save-path (gates the POST /slots endpoint regardless of whether save/restore is used) - Add requests to requirements.txt Permission pre-grant - Add permission block to tools/opencode/opencode.json granting allow for read, glob, grep, edit, write, and external_directory - Eliminates the interactive permission dialog on every invocation: OpenCode raises external_directory for all workspace paths because the ephemeral workspace has no .git root, causing project ID to resolve to "global" (anomalyco/opencode#8538) - Permission state is in-memory only and not persisted across processes (confirmed in permission/next.ts), so dialog would fire on every run regardless of prior approvals - Permission approval fallback (Tab+Enter for Allow always) retained in run_opencode_tui with saw_post_submit_output reset after each approval to prevent READY_RE match on post-approval redraw from prematurely terminating the response-wait loop Regex output assertions - Add assert_matches_all and assert_not_matches_any functions accepting regex patterns, wired to new YAML keys stdout_contains_re, stdout_not_contains_re, stderr_contains_re, stderr_not_contains_re - Required because TUI terminal wrapping splits long paths across lines, breaking literal stdout_contains assertions for nested paths - Update glob_discovery.yaml to use stdout_contains_re with pattern src[\s/]*nested[\s/]*sample\.py, matching the path regardless of how the TUI wraps it across line boundaries Other - Increase quiet_timeout from 2s to 5s to allow TUI rendering to complete before the harness declares the response done - Update write_file.yaml prompt to explicitly invoke the write tool, preventing the model from responding with a plan instead of executing - Add requests and pexpect to README install instructions
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
feat(infra): llama-server VRAM analysis, GGUF metadata inspection, and optimized server configuration
Overview
Three new tools and an updated server launch script, developed through
iterative empirical analysis of gpt-oss-20b MXFP4 running under
llama-server on an RTX 4090. The work uncovered non-obvious behavior
around KV cache allocation, slot concurrency, and context length
management that informed the final server configuration.
llama_server_vram_report.py
A log file parser that extracts and pretty-prints a structured VRAM
metrics report from a llama-server startup log.
Sections reported:
length, KV quantization
SWA window size, embedding dimension, attention heads (Q and KV),
head dimensions, MoE expert count and active experts per token,
expert utilization percentage, RoPE scaling type and factor
breakdown (model weights CUDA, CPU mapped, KV cache non-SWA, KV
cache SWA, compute buffer), total used, free after load
(with capping warning when --ctx-size exceeds model training
context), non-SWA and SWA cell counts and layer counts, per-slot
KV cost
margin, usable headroom, additional slots possible
Key findings documented through empirical log analysis:
KV cache is O(ctx_size), compute buffer is also O(ctx_size).
The compute buffer grows approximately linearly with --ctx-size at
~768 MiB per 131,072 tokens, making it the binding constraint when
scaling context — not the KV cache itself. The OOM boundary for
gpt-oss-20b MXFP4 on RTX 4090 lies between --ctx-size 524,288 and
655,360 at --parallel 16 --kv-unified.
kv_unified = true is not the default when --parallel is set
explicitly. When --parallel is omitted, the server auto-selects
n_parallel = 4 and sets kv_unified = true (hardcoded in server.cpp).
When --parallel N is passed explicitly, kv_unified defaults to false,
causing the pool to be divided equally across slots (ctx_size /
n_parallel tokens per slot). The --kv-unified flag must be passed
explicitly alongside --parallel to preserve full per-slot context.
In unified mode, slot count is essentially free. The KV pool size
is fixed by --ctx-size regardless of --parallel. Adding slots from 4
to 8 to 16 produced identical total KV allocation (1,632 MiB) with
negligible overhead. The per-slot cost figure in the capacity analysis
is therefore misleading in unified mode and should be interpreted as
a scheduling budget, not a memory budget.
--ctx-size is capped per slot at the model's training context.
Setting --ctx-size 524,288 with --parallel 4 does not give 4 slots
at 512K context. llama-server caps each slot at the model's training
context length (131,072 for gpt-oss-20b), logging a warning:
"the slot context (524288) exceeds the training context of the model
(131072) - capping". The pool holds 524,288 cells total but each slot
is limited to 131,072. The report surfaces this via a "Context per
slot" row with a capping note.
Bug fixes during development:
initializing slots, n_slots = Nfirst(present when --parallel is explicit), then the auto-detection
format, then the bare n_parallel fallback. Previously only the
auto-detection format was checked, causing explicit --parallel
invocations to report 1 slot.
(fit probe + actual load) to avoid double-counting.
fit probe's 0.00 MiB entry.
gguf_model_info.py
A standalone GGUF metadata inspector using the gguf and prettytable
packages. Reads architecture metadata directly from the model file
without requiring a running server.
Sections reported in Unicode box-drawing tables:
dimension, feed-forward length, vocabulary size
window size and pattern, RMS norm epsilon
utilization percentage, expert groups, shared experts, expert
FF length
scaling type and factor, original context length, context
extension multiplier
quantization-type breakdown with tensor count and parameter
count with percentage of total
The tensor breakdown for gpt-oss-20b MXFP4 confirms the naming:
MXFP4 accounts for 91.4% of parameters (72 tensors, expert weight
matrices), Q8_0 accounts for 8.6% (98 tensors, attention weights),
and F32 accounts for 0.04% (289 tensors, norm weights and embeddings).
Despite the majority of tensors being F32, they represent a negligible
fraction of total parameters.
Uses Keys constants from the gguf package for field name resolution,
making the script architecture-agnostic — the {arch} template is
resolved at runtime from general.architecture.
Requires: pip install gguf prettytable
llama-server-openai_gpt-oss-20b-MXFP4.sh (updated)
Updated server launch script incorporating findings from the VRAM
analysis work.
Changes from prior version:
when --parallel is specified explicitly
the relationship between slot count and pool size explicit and
self-documenting
halving KV cache VRAM vs f16 (408 MiB/slot vs 768 MiB/slot)
POST /slots/{id}?action=erase endpoint used by the pytest harness
KV cache flush fixture
under a shared TMP=/tmp/llama-server root
$(llama-server --version 2>&1 | tr -s ' \n' ' ' | xargs)
set via system prompt "Reasoning: ") and the slot/KV cache
architecture informed by source analysis
script for easy reconfiguration
Resultant configuration on RTX 4090 with gpt-oss-20b MXFP4:
llama_server_slots_brief_r1.md
A technical reference document covering the slot system architecture,
KV cache allocation modes, the concurrency model and its distinction
from CPU thread parallelism, the KV cache flush API, and the
relationship to OpenCode's parallel agent roadmap. Intended as a
reference for contributors working with llama-server in multi-session
or multi-agent contexts.