Proudly Made in Nebraska. Go Big Red! 🌽 https://xkcd.com/2347/
📝 Background reading: Mismanaged Geniuses — why recursive language models matter, from Alex Zhang's interview on RLMs.
An open-source Recursive Language Model (RLM) harness, written for the Qwen model
served by Strata on pluto. Apache-2.0.
A Recursive Language Model treats a long input as an environment rather than a
prompt. Instead of stuffing a million-token document into the context window, the
harness loads it into a persistent Python REPL as a variable (context) and asks the
model to write code that inspects, slices, searches and summarises it. The model only
ever sees the short outputs of the code it runs.
Inside that REPL the model can call the language model again (llm_query) on a chunk
of the data, or start a nested RLM (rlm_query) that gets its own REPL. Results stay in
variables; only small summaries flow back up. The run ends when the model marks an
answer as final. This is the scheme described in Recursive Language Models
(Zhang, Kraska and Khattab, arXiv 2512.24601),
which shows a frontier model handling inputs far beyond its native context window and
degrading gracefully as inputs grow.
ReCLamO-Harness is a clean-room implementation of that loop, sized like
rlm-minimal and borrowing the useful parts of the full rlm package (depth
recursion, batched sub-calls, limits, JSONL trajectories). What is new here is the
Qwen tuning: <think> handling, the paper's "do not over-call" guidance, hard caps on
sub-calls, and a prompt sized for the 32K of KV cache that stays resident on pluto.
Evaluation test plan: see
evals/TEST-PLAN.mdfor what is measured, in what order, and how results are judged.
Pre-alpha, usable. Work is tracked in epic
#8 and mirrored in
BACKLOG.md. reclamo ping and reclamo run work end to end against
pluto; the Docker sandbox and the live eval are the remaining children.
uv sync
uv run reclamo ping --profile pluto
uv run reclamo run --profile pluto --context ./big.txt -q "Which line holds the needle?"ping lists the server's models, sends one short non-thinking request and one
thinking request, and reports latency, token usage and whether reasoning came back.
run loads the context into a REPL and drives the model until it answers:
| Flag | Meaning |
|---|---|
--context PATH|- |
a text file; a directory (loaded as {relative path: text}, hidden and non-UTF-8 files skipped); or - for stdin |
-q, --query |
the question |
--max-depth N |
recursion depth; default 1 (the paper found Qwen gets worse at depth 2+; see Depth-2 recursion under Results) |
--max-iterations N |
root turns before a forced finish; default 20 |
--log-dir DIR |
where trajectories go; default runs/ |
--verbose |
print each turn (response, code, output) to stderr |
--no-thinking |
disable thinking on root turns |
--json |
print the full result (answer, turns, sub-calls, usage, stop reason, trajectory path) as JSON |
--sft |
also write the SFT file (see below) |
--sandbox {subprocess,docker} |
where generated code runs (default subprocess) |
--protocol {fence,tools} |
how the root model acts: fenced code + FINAL text (default), or execute_python / final_answer tool calls (see Protocol under Results) |
Exit codes: 0 an answer was produced (including a forced finish when turns or time
run out: one last call that shows the model what the REPL holds and asks for FINAL /
FINAL_VAR, falling back to the answer dict, then an answer variable, then the reply's
prose with code removed), 3 a
limit stopped the run (token budget: any partial answer is printed; or max_errors
consecutive REPL errors, which since #50
also ends in the forced finish and prints its answer, stop reason error_limit), 2 configuration or key errors, 1 the endpoint failed.
max_timeout is a wall-clock bound
(#51): every root and sub-call
request is capped by the time left, no sub-call starts after the deadline, the rest of a
turn's blocks are skipped once it passes, and every forced finish is a single attempt
capped at the time left plus FORCED_FINISH_TIMEOUT (60 s). A run therefore ends within
max_timeout + 60 s + about 1 s (the shortest request timeout). The one exception is a
REPL block that is already computing at the deadline: it may finish its exec_timeout,
because killing the worker would lose the variables the forced finish reads.
Three profiles are built in.
plutopoints athttp://pluto:8083/v1, modelqwen3.8-flash-next, with Qwen's recommended sampling for thinking (root) and non-thinking (sub-call) turns, andcontext_tokens = 32768. That is the old resident window, and it keeps every request in Strata's fast range.pluto-longis the same server, model, key and sampling withcontext_tokens = 131072and a 30-minute request timeout for root, sub and the plain baseline. Seepluto-longbelow.openai-compatibleis a generic profile for any OpenAI-style server.
Concurrency 2 on pluto (since 2026-10-08). Strata now serves two requests at once
("parallel": 2, int8 KV, 64K resident; /v1/status reports concurrency.serving=2).
pluto and pluto-long therefore set concurrency = 2, so llm_query_batched uses
both slots. Measured on that change: single-request decode fell from 46 to 35 tok/s,
and a 63K-token prompt read at 227 tok/s. Concurrency is per endpoint. A planner on its
own base_url keeps its own concurrency (the two-endpoint examples below keep the
planner at 1). To undo it, set concurrency = 1 on the profile. Runs before
2026-10-08 used 1.
Sub-call thinking is a profile switch, [profiles.<name>.sub] enable_thinking, or
--sub-thinking for one reclamo run. It is off on pluto because Strata is slow and
serial per slot. The paper's sub-model was a reasoning model (GPT-5 with GPT-5-mini
sub-calls), so turning it on is the closer match when time allows.
Add or override profiles in ~/.config/reclamo/profiles.toml (or the file named by
RECLAMO_PROFILES). Fields you leave out keep the built-in values:
[profiles.pluto]
max_iterations = 30 # override one field of a built-in profile
[profiles.lab] # a new profile
base_url = "http://lab:8000/v1"
concurrency = 2
context_tokens = 65536
subcall_chars = 16000 # per-call size the prompt advertises (default: derived, see below)
prompt_version = "v0.2" # root prompt version: v0.2 (default) or v0.1 (published results)
max_subcalls_per_run = 64
max_subcalls_per_exec = 24
exec_timeout = 120.0 # seconds per REPL execution
max_timeout = 1800.0 # seconds for the whole run
[profiles.lab.root] # the loop's own turns
model = "qwen3-32b"
max_tokens = 4096
enable_thinking = true
reasoning_effort = "medium"
sampling = { temperature = 0.6, top_p = 0.95, top_k = 20, min_p = 0.0 }
[profiles.lab.sub] # llm_query calls
model = "qwen3-32b"
max_tokens = 2048
enable_thinking = false
sampling = { temperature = 0.7, top_p = 0.8, top_k = 20, presence_penalty = 1.0 }The root (planner) and sub (reader) roles can talk to different servers. Each role
table may set base_url, api_key_env, api_key_cmd and concurrency; anything it
leaves out inherits the profile-level value, so existing profiles are unchanged. Each
distinct base_url gets its own HTTP client, key and semaphore: sub-calls to pluto
are not queued behind root turns on another host, while roles that share a URL still
share one semaphore (pluto's two slots are shared by its roles). Roles that share a URL must
agree on the key source and concurrency. reclamo ping checks every endpoint in the
profile and reports each one under its own base_url: heading.
This profile is not built in. It shows the shape for a locally hosted Qwen3-8B planner (epic #42) driving the harness, with pluto's Flash-Next answering the sub-calls. The planner model name is a placeholder until the trained adapter exists:
[profiles.planner-8b] # profile level = the reader (pluto), as in [profiles.pluto]
base_url = "http://pluto:8083/v1"
api_key_cmd = "pass pluto/flashnext-api-key"
concurrency = 2 # pluto's two slots; the planner below keeps its own 1
context_tokens = 32768
[profiles.planner-8b.root] # the planner, on a local server
base_url = "http://localhost:8080/v1"
api_key_env = "RECLAMO_PLANNER_API_KEY" # any value if the local server ignores keys
concurrency = 1
model = "qwen3-8b-rlm" # placeholder
max_tokens = 4096
enable_thinking = true
reasoning_effort = "medium"
sampling = { temperature = 0.6, top_p = 0.95, top_k = 20, min_p = 0.0 }
[profiles.planner-8b.sub] # the reader: inherits pluto's URL, key and concurrency
model = "qwen3.8-flash-next"
max_tokens = 2048
enable_thinking = false
sampling = { temperature = 0.7, top_p = 0.8, top_k = 20, presence_penalty = 1.0 }Environment overrides apply to any profile: RECLAMO_BASE_URL (one URL for every
role; it replaces per-role base_urls too), RECLAMO_MODEL (sets both roles) and
RECLAMO_API_KEY. When RECLAMO_API_KEY is unset the key
comes from the profile's api_key_cmd (pass pluto/flashnext-api-key for pluto).
Non-standard sampling keys such as top_k and min_p are sent in the request body,
and enable_thinking goes out as chat_template_kwargs, which Strata and vLLM-style
servers honour.
The paper's authors released their fine-tuned planner,
mit-oasys/rlm-qwen3-8b-v0.1
(MIT; uploaded 2026-01-15). The model card says it "was trained on trajectories produced
using a fixed system prompt" and "assumes the environment/scaffold from our RLM repo".
planner_style = "upstream-rlm-v0" makes the root speak that format, so the checkpoint can
plan here without retraining (issue #43,
epic #42 decision 7). Sub-calls
still go to the reader endpoint, and our sandbox, sub-call caps, nudges, compaction, forced
finish and logging all stay on. The default style stays reclamo.
What the style reproduces, and where it comes from:
| Piece | Upstream format | Source |
|---|---|---|
| System prompt | the paper's Qwen3-8B / 32K prompt: the GPT-5 prompt (1a) with the (1c) diff applied (32K warning, llm_query "~100k chars", batching rule, context[:1000], a line-start FINAL_VAR(final_answer) example, "NOT in code or repl tags"); no llm_query_batched |
arXiv 2512.24601 v2 (2026-01-28) App. C.1 |
| Message order | system, then the context metadata as an assistant message (Your context is a str with N total characters, and is broken up into chunks of char lengths: [...]), then per turn the assistant reply and one user message per executed block |
alexzhang13/rlm v1.0.0 (18a9368, 2026-01-12), rlm/utils/prompts.py, rlm/core/rlm.py |
| Per-turn prompt | USER_PROMPT_WITH_ROOT with the query, plus the "You have not interacted with the REPL environment..." safeguard on turn 1 and "The history before is your previous interactions..." after; sent every turn, never stored |
same, build_user_prompt |
| Code | ```repl fences (we also accept ```python) |
same, rlm/utils/parsing.py |
| REPL output | Code executed:\n```python\n<code>\n```\n\nREPL output:\n + stdout, stderr and REPL variables: [...], cut at 20,000 chars with ... + [N chars...] |
same, format_iteration, format_execution_result |
| Finish | FINAL(...) / FINAL_VAR(name) at the start of a line; a FINAL_VAR in the same reply as code is read after the code runs, as upstream does |
same, find_final_answer |
The prompt text is copied verbatim (with upstream's copyright and MIT notice beside it) in
src/reclamo/upstream.py. What is adapted: the variable list shows context and our user
variables (our REPL has no context_0); an exception's line carries its line number; our
nudges (a rejected FINAL, no action, re-verify, decompose) arrive as their own user message
after the outputs; the forced finish restates the query, which this style never stores; and
the answer dict is switched off so answer = llm_query(...), as in the prompt's own
example, is an ordinary variable. The tools protocol is not available with this style.
This profile is not built in. The planner runs on a local llama.cpp server with the community
GGUF cameronbergh/rlm-qwen3-8b-v0.1-gguf
(q8_0 or f16), for example llama-server -m rlm-qwen3-8b-v0.1-q8_0.gguf --jinja -c 32768 --port 8080; the reader is pluto. The URL and model name are placeholders:
[profiles.rlm-qwen3-8b] # profile level = the reader (pluto)
base_url = "http://pluto:8083/v1"
api_key_cmd = "pass pluto/flashnext-api-key"
concurrency = 2 # pluto's two slots; the planner below keeps its own 1
context_tokens = 32768
planner_style = "upstream-rlm-v0"
output_truncate_chars = 20000 # upstream cuts each block's output at 20,000 chars
[profiles.rlm-qwen3-8b.root] # the paper's planner, on a local llama.cpp server
base_url = "http://localhost:8080/v1"
api_key_env = "RECLAMO_PLANNER_API_KEY" # any value if the local server ignores keys
concurrency = 1
model = "rlm-qwen3-8b-v0.1" # placeholder
max_tokens = 4096
enable_thinking = false
sampling = { temperature = 0.7, top_p = 0.8, top_k = 20, min_p = 0.0 }
[profiles.rlm-qwen3-8b.sub] # the reader: inherits pluto's URL, key and concurrency
model = "qwen3.8-flash-next"
max_tokens = 2048
enable_thinking = false
sampling = { temperature = 0.7, top_p = 0.8, top_k = 20, presence_penalty = 1.0 }reclamo run --profile rlm-qwen3-8b ... uses it; --planner-style upstream-rlm-v0 selects
the style on any profile.
Sampling and thinking: the model card and generation_config.json give no sampling values.
The checkpoint ships a Qwen3-Coder-style chat template with no thinking switch, and it was
distilled from Qwen3-Coder-480B-A35B-Instruct, a non-thinking model, so thinking is off and the
sampling is Qwen3's non-thinking preset (temperature 0.7, top_p 0.8, top_k 20, min_p 0).
Caveats:
- The exact training prompt is not public. The checkpoint's training code was never released
(the repo's
training/harness, added 2026-05-24, is for the later 30B model and the newer answer-dict format). The paper says the 8B experiment used the (1c) prompt; its listing puts the context-metadata line inside the system prompt, while the repo sends it as an assistant message. This style follows the repo, which the model card points to. Where the (1c) diff inserts the bareFINAL_VAR(final_answer)line is reconstructed from the hunk offsets. - The trajectories were Qwen3-Coder-480B's, sampled on LongBenchPro, with sub-calls answered by Qwen3-8B. Here the reader is Flash-Next, whose answers will read differently.
- The prompt advertises
llm_queryat "~100k chars". Our per-block and per-run sub-call caps still apply, and a block that exceeds them gets an error, which the checkpoint never saw. - 20,000-character outputs are large for a 32K window, and this checkpoint writes about ten
blocks a turn. Compaction keeps every request inside the window (see Context budget below),
shrinking even the newest turn's outputs when it must. Lower
output_truncate_charsso the planner sees more of each output. - Not yet run against a hosted copy: that smoke run is #40.
#59. prompt_version
selects the reclamo root prompt (fence and tools protocols alike). planner_style = "upstream-rlm-v0" ignores it and stays verbatim.
v0.2(default) follows the paper's main prompt and upstream's current prompt on delegation.v0.1is the prompt every result in this README was measured with, kept byte for byte (system prompts, nudge text and the 12,000-character per-call hint). Use--prompt-version v0.1(onreclamo run,bench.pyandrun_eval.py) orprompt_version = "v0.1"in a profile to reproduce them.
What v0.2 changes, and why:
| v0.1 | v0.2 | Source | |
|---|---|---|---|
| Stance on sub-calls | "IMPORTANT: sub-calls are expensive … never 1000 … further calls fail with an error" | "you are strongly encouraged to use them as much as possible … don't be afraid to put a lot of context into one call". The hard caps are stated as plain facts | Paper App. C.1 (1a), the main prompt |
| Planning | none | "Work as an orchestrator, not a solver": probe, state the decomposition and the sequence of turns, then delegate. It also says when not to delegate: if a keyword/regex search already pins the answer, or one visible passage holds it, read it directly | Upstream ORCHESTRATOR_ADDENDUM (rlm/utils/prompts.py, added in de762b9 on 2026-05-24, on by default) |
| Examples | chunk-and-map, split-on-structure | one call when the data fits; map then aggregate with a final llm_query; a running buffer carried from call to call |
Paper App. C.1 (1a): the book example (whose undefined buffers is fixed here) and the per-chunk "aggregating all the answers" call |
| Per-call size | fixed 12,000 characters | derived from the sub model's window: (window − sub.max_tokens) × 0.85 × 3 characters (78,336 on pluto); subcall_chars still overrides it, and a sub endpoint can set its own context_tokens |
Paper (1a): "Analyze your input data and see if it is sufficient to just fit it in a few sub-LLM calls!" |
| Decompose nudge | fires when one call gets ≥90% of a context over 12,000 characters | the same, but measured against the derived size, so it never fires when the whole context fits in one sub-call. Its text gives both sizes | the same |
Why the warning went. The paper's GPT-5 prompt (1a) has no warning. Its "IMPORTANT: Be
very careful about using llm_query as it incurs high runtime costs" line was added for
Qwen3-Coder (1b) and the Qwen3-8B variant (1d), because "without this warning, the model
will try to perform a subcall on everything, leading to thousands of LM subcalls for
basic tasks" (App. C). Our failure is the opposite: in the #22 re-run only 7 of 40 rlm
rows made any sub-call. The caps (max_subcalls_per_exec, max_subcalls_per_run) still
enforce the limit, so the prompt does not need to.
The size formula is conservative on purpose. The sub-call prompt also carries the question. Digit-heavy text runs close to one token per character, while prose runs near four. An oversized prompt fails as a server error inside the REPL, which the model sees and can recover from.
The copied upstream text keeps its MIT notice beside it in src/reclamo/prompts.py.
The prompt text is not the only thing v0.2 changes: see the defaults below. Sampling is
unchanged.
#61. These defaults follow
prompt_version. A field you leave unset takes its version's value, and a value set in
a profile, the user file or code always wins. prompt_version = "v0.1" restores every
old value, so published results stay reproducible. (One caveat: under v0.1 an unset sub
max_tokens is now 2,048, the built-in profiles' old value, where a hand-written
profile used to get 4,096.)
| Setting | v0.1 | v0.2 | Why / source |
|---|---|---|---|
max_iterations |
20 | 30 | Upstream rlm main default (rlm/core/rlm.py, _DEFAULT_MAX_ITERATIONS = 30 in rlm/utils/prompts.py). In the #22 re-run, 15 of 40 rlm rows used all 20 turns |
output_truncate_chars |
2,000 | 20,000 | Upstream v1.0.0 rlm/utils/parsing.py cuts each REPL output at 20,000 chars; its prompt says "REPL outputs over ~20K characters are truncated". Compaction (#49) shrinks oversized turns to fit the window (tested at 20K on a 32K window) |
root max_tokens |
4,096 | 8,192 with thinking on (4,096 off) | Paper App. B: "Thinking models without sufficient output tokens struggle as RLMs" |
sub max_tokens |
2,048 | 4,096 | Room for a full answer from a big sub-call |
max_subcalls_per_run |
64 | 256 | The paper's failure mode was "thousands of LM subcalls" (App. C), not hundreds. The cap stays as a safety net |
max_subcalls_per_exec |
24 | 64 | The same |
root / sub timeout (per request) |
300 s | 900 s | An 8,192-token thinking reply at ~35 tok/s (Strata, two slots) is ~234 s of decode before any prompt reading; at 300 s long root turns timed out and retried. pluto-long sets 1,800 s |
max_timeout (also bench.py / run_eval.py) |
1,800 s (scripts: 600 / 900 s) | 3,600 s | Time is not the constraint (CJ, 2026-10-08) |
| reply cut off by the output limit | retried without thinking only if empty | also continued once (thinking off) when cut part-way with no usable action, and retried without thinking when the "content" is cut-off reasoning | Paper App. B, the same finding |
Unchanged on purpose: max_depth (1), protocol (fence), sampling, and pluto's
context_tokens (32,768).
A cost on pluto: the larger root reply reserve shrinks the plain baseline's usable
window from 28,672 to 24,576 estimated tokens. GLaDOS small (about 26.8K estimated
tokens) no longer fits there. Use pluto-long for plain.
Strata accepts prompts far beyond its resident KV (--max-context 262144, with
--kv int8 --kv-resident 65536; a 62,897-token prompt read at 227 tok/s on
2026-10-08), but on pluto the profile's
context_tokens = 32768 also capped the plain baseline, whose usable window is
context_tokens minus the root output reserve. Every medium and large plain row was
therefore "does not fit". That was our limit, not the hardware's.
pluto-long sets context_tokens = 131072. That is not Strata's full 262K: medium
contexts run about 90–150K estimated tokens, and the estimate runs high on
digit-heavy text. It also sets a 1,800 s request timeout for root and sub (150K tokens
at ~200 tok/s is about 750 s of prompt reading before generation). The plain call uses
the root timeout, so it is covered too. Everything else matches pluto.
With --profile pluto-long the plain baseline fits all 16 small and 14 of 16 medium
cells of the independent eval (seeds 0 and 1). Only GLaDOS medium, at about 148K
estimated tokens, still does not fit. No large cell fits. The cost is speed: prompt
reading drops from ~500 to ~200 tok/s past the resident window, so a 100K-token plain
call spends about 8 minutes reading. Under v0.2 the advertised sub-call size follows
the window too, about 324K chars, which makes "one big sub-call" possible at the same
cost.
Every root request is kept under an input budget of context_tokens − the root role's
max_tokens − 10% of context_tokens (the margin absorbs estimate error), and never less
than a quarter of context_tokens (#49).
The estimate counts one token per digit and 3.5 characters per token otherwise, and it
includes the tool specs and upstream's per-turn prompt. Before each turn the history is
brought under the budget in three steps:
- REPL outputs older than the last four are replaced by stubs and, if that is not enough, summarized with one sub-call (unchanged).
- If it still does not fit, as when one turn's outputs alone overflow the window, the
remaining REPL outputs are shrunk, largest first and oldest first among equals. Each
keeps its head and tail around a stub that says how much was cut and that the full
text is in the REPL's
historyvariable. The model's own code, the notes and the turn line are never cut. - If the request still cannot fit, it is not sent. The run logs
context_overflowand takes the forced finish (stop reasoncontext_overflow) with only the opening messages and the REPL state. If even that is over the budget, there is no model call and the run falls back to the REPL values.
Every run writes runs/<UTC timestamp>-<id>.jsonl, one JSON object per line:
metadata— profile (never the key), query, context shape, depth, andendpoints(each distinctbase_url, the roles it serves and its concurrency)iteration— turn number, the model's content, its reasoning as a separate field, the code blocks, each block's result, any nudges, the FINAL decision, and theendpointandmodelthat served the turnsubcall— depth, kind (llm_query,llm_query_batched,rlm_query), prompt and answer sizes, latency, running count, and theendpointandmodelthat answered (absent for anrlm_querythat recursed; the child's own turns carry them)compaction— when REPL outputs were elided, summarized or shrunk to fit (sizes, budget,elided_turns,shrunk_messages)context_overflow— a request could not be made to fit the window: the estimated size, the budget,context_tokensandmax_tokens, and what the forced finish did insteadforced_finishandfinal— how the run ended.forced_finish.sourcesays where a forced answer came from:model(a validFINAL/FINAL_VAR/final_answerin the forced reply),answer_dict,variable(with its name),reply_text(the reply's prose, code and thinking removed) orempty
Reasoning is logged but never fed back into the history, per Qwen's guidance. With
--sft a second file, <same stem>.sft.jsonl, holds one line per root turn,
{"messages": <history before the turn>, "completion": <content>}, the format the
paper used to fine-tune RLM-Qwen3-8B.
The model writes code, and that code runs. Two things keep that contained.
Subprocess REPL (default). Generated code runs in a separate python -I process
whose environment is scrubbed to PATH, LANG, LC_ALL and a throwaway HOME and
TMPDIR. No RECLAMO_*, OPENAI_* or *_API_KEY variable reaches it. The worker's
protocol stdio is duplicated to private descriptors before user code runs, and fd 0
is pointed at /dev/null and fd 1 at stderr, so generated code (or anything it
spawns) can neither read protocol messages nor corrupt the stream. Each execution
has a wall-clock timeout; a hung or crashed worker is killed and relaunched, and the
model is told its variables are gone. Sub-calls per execution and per run are capped
in code, not just in the prompt.
Docker sandbox (--sandbox docker, #6).
The same worker runs in python:3.12-slim with --network none, a read-only root
filesystem, a small /tmp, no capabilities and a non-root user. The only channel out
is the JSON-lines protocol on stdio, which is how llm_query still works with no
network.
What the key can reach. The API key lives only in the parent process and is used for one thing: requests to the configured LM endpoint. It is never passed to the REPL, never written to a trajectory, and never included in an error message.
-
examples/needle.py— builds N synthetic lines (default 1,000,000, about 30 MB) with one passphrase line and asks for the passphrase. Exit 0 when found. -
examples/oolong_lite.py— about 300 synthetic support tickets in five categories, written as paraphrases that never contain their category's words (asserted at generation time), so grep cannot solve it and the model has to read, which means batching tickets intollm_querycalls. Scored by per-category absolute error.--dumpprints the tickets and the ground truth without calling a model. -
examples/eval.py— runs both against a profile, writesruns/eval-<timestamp>.jsonand prints the markdown table used below. -
examples/nested.py— the depth-2 task: a dict of 12 department ticket logs (~25K chars each, four different line formats) where each department needs format discovery, an open/close/reopen replay in code and a semantic read of the surviving complaints (the OOLONG-lite paraphrases, so grep cannot classify them). Scored by per-department absolute error plus whether the "most" department is right.--max-depth,--delegate(tell the model to userlm_queryper department),--max-timeout,--dump,--json. -
examples/longdoc_qa.py— two-hop questions over a long synthetic engineering digest: each project's owner and that person's office city sit in different sections among distractors, so answering needs two hops of reading. Exact-match on the city. -
examples/bench.py— the repeatable benchmark: harness vs plain model, N seeds over a task × size grid, resumable (--resume), with a markdown summary (--summary-only). -
evals/independent/run_eval.py— the frozen independent eval (#22): the same rlm and plain modes on the eight roundtable-authored tasks, scored by each task's ownscore(); resumable, with a markdown summary. -
evals/longbenchpro/longbenchpro.py— the LongBench Pro loader (#75), described below.
uv run python examples/needle.py --profile pluto --lines 1000000
uv run python examples/oolong_lite.py --profile pluto --tickets 300
uv run python examples/eval.py --profile pluto
uv run python examples/bench.py --profile pluto --seeds 3
uv run python examples/nested.py --profile pluto --max-depth 2 --delegate --max-timeout 900Live tests (uv run pytest -m live) run ping, a thinking round trip and a
50K-line needle against pluto; they are skipped by default and in CI.
caskcsg/LongBench-Pro
(Apache-2.0, arXiv 2601.02872) has 1,500 items on real documents: 11 primary and 25
secondary tasks, half English and half Chinese, and six length buckets from 8k to 256k
Qwen tokens. It ships as one test split. The loader cuts it into three splits for the
trained-planner epic (#42):
| Split | Use | Items |
|---|---|---|
practice |
training data | 1,110 |
practice-dev |
recipe and checkpoint selection | 126 |
held-out |
evaluation only | 264 |
- Splits are fixed. No random seed is involved: the split comes from a salted sha256 of an item id. The 1,500 items have only 1,076 distinct contexts, and one source text often recurs across length buckets. Items whose contexts are equal, or share a normalised line of at least 120 characters, form one document group (577 in all). The whole group lands in one split, keyed by its smallest id, so no document is split across practice and held-out.
- MC vs open is tagged per item. An item is
mcwhen its answer is one option letter that the question lists, andmc-multiwhen the question lets the model pick several letters (any non-empty subset). Everything else isopen. The totals are 203mc, 103mc-multiand 1,194open. T3 and T11 are MC, but 83 of their 240 items (20 in T3, 63 in T11) are multi-select, which a random guess almost never passes, and T11 has one open item. T10 also has 62 MC items. Every item also carrieschance, the exact-match rate of a uniform guess. - The exam is excluded.
exam_hashes.jsonholds the sha256 of every activeevals/independentdocument (8 tasks, 3 sizes, seeds 0-99), after NFKC, casefolding and whitespace collapse. Any item whose context or question matches is refused, whatever its split. None match today, as expected for synthetic exam documents. A test fails if an exam generator changes and the hashes were not rebuilt. - Scoring follows the upstream metrics, reimplemented here: first-line accuracy,
SubEM, F1, pairwise order and NDCG, chosen per secondary task. T4 summaries need an
embedding model, so they are left unscored and filtered out by default.
instance(item)returns abench.Instancewhose task is registered inbench.TASKS, sobench.run_rlmandbench.run_plainrun it unchanged.
The data (530 MB) is downloaded with the standard library at a pinned revision and checked against its sha256. No new dependency is needed. Tests use a small synthetic fixture and never touch the network.
uv run python evals/longbenchpro/longbenchpro.py download # to ~/.cache/reclamo/longbenchpro/
uv run python evals/longbenchpro/longbenchpro.py stats # per task x kind x split
uv run python evals/longbenchpro/longbenchpro.py stats --split held-out --max-bucket 16k
uv run python evals/longbenchpro/longbenchpro.py exam-hashes # after any exam generator changestats reports counts, median characters, and the median, p90 and max token estimate
(bench.estimate_tokens). The estimate overcounts English and undercounts Chinese, so
it also counts items by the dataset's own Qwen-token bucket. For a 16K-32K student
window, the held-out split has 82 scorable items in the 8k and 16k buckets, and 121 up
to 32k.
There are two kinds of result here. The independent evaluation comes first: it uses tasks written by eight non-Anthropic models that had no part in building the harness. The smoke, benchmark, protocol and depth-2 results after it use tasks written by the same Claude models that wrote the harness and its prompt, so they likely overstate what the harness can do.
Correction (issue #57). Every plain number published in this README before #57 is best-of-2 selected on the answer key: each plain row made two calls (thinking on, thinking off), scored both against the truth and kept the better one, while the harness got one attempt. That inflates plain. On the #22 re-run's 16 small cells, plain was exact on 12 as best-of-2, but on 9 with thinking on alone (the pluto root setting) and on 8 with thinking off alone. The tables below keep the numbers as they were measured. Since #57,
plainis one call with thinking as the profile's root role sets it (the same model configuration as the harness's root), andplain-think/plain-nothinkare explicit single-call modes, each its own row.--summary-only FILE --legacy-plain thinking(ornothink) re-summarises an old row file as one variant.
#22. The tasks are the eight
generators in evals/independent/fixed/active/, written by the
non-Anthropic lanes of the Flatline Roundtable. Each answer is scored by that
generator's own score(), where 1.0 means exact. We did not write or change any
question, answer key or scorer.
Frozen harness: commit e17580fb76b06ece39ba381ce26763aa870bca90. Nothing was
tuned on these tasks. The run used the default pluto profile, the fence protocol,
depth 1 and default caps (20 turns, 64 sub-calls), plus max_timeout=900 s per run.
It took 3.5 hours on 2026-10-07, one row at a time under the pluto lock. Every
batch's metadata records that src/, examples/ and pyproject.toml were unchanged
from the frozen commit.
The model was the same as in the sections below: huihui-qwen3.8-flash-next-abliterated
served by Strata on pluto. The runner is
evals/independent/run_eval.py, which imports
bench.py's two modes:
- rlm is the harness.
- plain is one call with the whole context in the prompt. In this run (before #57) it was the better of thinking on (16K output cap) and thinking off, chosen against the answer key; see the correction above. It is recorded as "does not fit" when the prompt is larger than pluto's 28,672 usable tokens.
The grid was 8 tasks × small (~60K chars), medium (~300K) and large (~1.2M), with seed
0 for every cell and seed 1 for small and medium: 80 rows and no errors. Raw rows are in
runs/independent-20261007.json, with one trajectory per rlm row in runs/ (not in
git).
Each cell shows seed 0 / seed 1. † marks a cell whose answer key cannot be reached from the context (see "Defects in the tasks" below).
| Task (author model) | Size | rlm score | plain | rlm turns | rlm sub-calls | rlm s | rlm stop |
|---|---|---|---|---|---|---|---|
| GLaDOS (grok-4.6) | small | 1 / 0.60 | 1 / 1 | 20 / 20 | 0 / 0 | 318 / 488 | out of turns / answer_dict |
| medium | 0.20 / 0.60 | does not fit (~148K tokens) | 20 / 20 | 0 / 0 | 296 / 247 | out of turns / out of turns | |
| large | 0.20 | does not fit (~606K tokens) | 20 | 0 | 193 | out of turns | |
| SHODAN (gpt-6-astra) | small | 0 / 0 | 1 / 0 | 20 / 20 | 10 / 0 | 475 / 178 | out of turns / out of turns |
| medium | 0 / 0 | does not fit (~88K tokens) | 20 / 20 | 2 / 0 | 275 / 134 | out of turns / out of turns | |
| large | 0 | does not fit (~353K tokens) | 20 | 0 | 168 | out of turns | |
| TheDixieFlatline (gemini-3.1-pro-high) | small | 0† / 0† | 0† / 0† | 12 / 11 | 0 / 0 | 108 / 231 | final / answer_dict |
| medium | 0† / 0 | does not fit (~96K tokens) | 14 / 17 | 0 / 0 | 99 / 181 | answer_dict / final | |
| large | 0 | does not fit (~396K tokens) | 20 | 0 | 180 | out of turns | |
| Cerebex (glm-5.3-flash) | small | 0 / 0 | 0.50 / 1 | 20 / 20 | 6 / 0 | 379 / 353 | out of turns / out of turns |
| medium | 0 / 1 | does not fit (~93K tokens) | 20 / 20 | 0 / 0 | 300 / 460 | final / out of turns | |
| large | 0 | does not fit (~373K tokens) | 20 | 0 | 242 | out of turns | |
| Neuromancer (deepseek-v4-flash) | small | 1 / 1 | 1 / 1 | 10 / 16 | 0 / 0 | 77 / 315 | final / answer_dict |
| medium | 0 / 1 | does not fit (~86K tokens) | 17 / 20 | 0 / 1 | 210 / 445 | final / final | |
| large | 0 | does not fit (~350K tokens) | 20 | 3 | 808 | out of turns | |
| SELMA (nemotron-3-super-120b-a12b) | small | 1 / 0† | 1 / 0† | 7 / 5 | 0 / 0 | 71 / 58 | answer_dict / final |
| medium | 1 / 0† | does not fit (~89K tokens) | 8 / 7 | 0 / 0 | 75 / 70 | final / answer_dict | |
| large | 0† | does not fit (~358K tokens) | 9 | 0 | 136 | final | |
| MasterControl (mistral-medium-3.1) | small | 0.50 / 1 | 1 / 1 | 11 / 12 | 0 / 0 | 209 / 106 | answer_dict / answer_dict |
| medium | 1 / 1 | does not fit (~86K tokens) | 14 / 9 | 0 / 0 | 211 / 113 | answer_dict / answer_dict | |
| large | 1 | does not fit (~343K tokens) | 9 | 0 | 124 | final | |
| Multivac (hermes-4-405b) | small | 0† / 0† | 0† / 0† | 7 / 4 | 0 / 0 | 54 / 23 | final / answer_dict |
| medium | 0† / 0† | does not fit (~104K tokens) | 5 / 5 | 0 / 0 | 67 / 43 | answer_dict / answer_dict | |
| large | 0† | does not fit (~428K tokens) | 5 | 0 | 29 | answer_dict |
What each task asks (from manifest.json):
- GLaDOS: the net authorized amount and certifying member of a renamed project's legal successor, under bylaws and an SOP.
- SHODAN: the earned credit per office after signed corrections, custody findings and agreement terms.
- TheDixieFlatline: the final room of an asset through renames and transfers.
- Cerebex: the travel total after amendments and reversals.
- Neuromancer: March travel reimbursed after adjustments and denials.
- SELMA: the final owner of a review after reassignments.
- MasterControl: the one employee flagged for two issues, and their total.
- Multivac: the status and description of a requirement after merges and splits.
Accuracy by size. A cell is one task × size × seed. A plain prompt that does not fit scores 0.
| Size | Cells | rlm mean | rlm exact | plain mean | plain exact | plain does not fit | Answerable cells | rlm mean | rlm exact | plain mean | plain exact |
|---|---|---|---|---|---|---|---|---|---|---|---|
| small (~60K) | 16 | 0.38 | 5/16 | 0.59 | 9/16 | 0/16 | 11 | 0.55 | 5/11 | 0.86 | 9/11 |
| medium (~300K) | 16 | 0.36 | 5/16 | 0.00 | 0/16 | 16/16 | 12 | 0.48 | 5/12 | 0.00 | 0/12 |
| large (~1.2M) | 8 | 0.15 | 1/8 | 0.00 | 0/8 | 8/8 | 6 | 0.20 | 1/6 | 0.00 | 0/6 |
Findings:
- Where the context fits, the plain model is clearly better. At ~60K characters, plain was exact on 9 of the 11 answerable cells; the harness managed 5. Plain also took a median 17 s for the chosen call, or 154 s for both variants, against 194 s for the harness. On SHODAN, Cerebex and GLaDOS small, plain was exact where the harness ran out of turns or finished with the wrong total. The self-authored benchmark below showed a tie at this size. These tasks show a loss.
- Where it doesn't fit, the harness is the only option, and it is weak. From ~300K characters up, plain cannot attempt any cell. The harness was exact on 5 of 12 answerable medium cells and 1 of 6 large ones, with a mean of 0.48 and then 0.20. It was reliable on one task, MasterControl (exact at medium and large on both seeds), and good on SELMA whenever the key was answerable. It won occasionally on Neuromancer, Cerebex and GLaDOS (partial credit). It never solved SHODAN at any size, nor TheDixieFlatline when that task was answerable. In the self-authored benchmark, by contrast, the harness was exact on every row where plain did not fit except one 1,000-ticket seed, up to ~16M tokens.
- The main failure is running out of turns. 15 of the 40 rlm rows hit the 20-turn
limit, and only 2 of those 15 were then exact. The trajectories show the same pattern
on the reconciliation tasks (SHODAN, Cerebex, GLaDOS): many turns printing sections
to read them, then a hand-written regex parser per section, then the limit.
One example is
runs/20261007T204104Z-a03836.jsonl, SHODAN small, where the forced finish says it is "out of turns but must answer" and guesses. - Qwen almost never delegates to sub-calls. Only 5 of the 40 rlm rows made any
llm_querycall. Each generator says it was built so that grep fails, yet the model greps and slices the context in code and then reads the slices itself. In the self-authored benchmark OOLONG-lite drew batched sub-calls. These tasks did not. - Clean protocol. 560 code executions produced 1 syntax error and 3 execution errors, with 0 rejected finals, 3 turns with neither code nor a final, and no row errors or stalls. The losses are in reasoning and the turn budget, not in the loop.
- Comparison with the earlier medium seed-0 run. The protocol section below scored the fence protocol at a mean of 0.28 with 2/8 exact on the same medium seed-0 cells. This run got 0.28 and 2/8 exact again, on the same two tasks. That run looked only at protocol and changed nothing; the harness is unchanged since.
Harness bugs found. Both were fixed after this run, in the forced-finish PR that
followed it (#22). The
numbers above are still the ones measured on frozen e17580f, and no row was re-run:
- The forced finish can return a stale REPL variable instead of the answer. When a
run runs out of turns,
_forced_finishreturns the first variable it finds namedfinal_answer,answer_text,resultorfinal. It does this before asking the model, and it does not check how recent the value is. In Neuromancer large (runs/20261007T220642Z-c69a18.jsonl), turn 20 printedFinal answer: $91049.09, which is exactly right. The run instead returnedresult, set on turn 18 to anllm_queryreply (a prose chain analysis), and scored 0.resultis a common name for a sub-call's reply, so this can recur. Fixed: the forced finish now always makes one more call. That call's prompt shows the current value ofanswer['content']and of each of those variables (300 characters each) plus the end of the last REPL output, and asks forFINAL(...)orFINAL_VAR(name). A valid reply wins. The variables are only a fallback, after the answer dict, so a good value already in the REPL is still kept (paper E.2) but no longer beats an answer the model just printed. - The forced finish can return code as the answer. When the model's forced-finish
reply has no
FINAL(...), the whole reply becomes the answer. Twice that reply was a```replblock, and the code was scored as the answer: SHODAN large seed 0 (runs/20261007T215541Z-947e9c.jsonl) and SHODAN small seed 1 (runs/20261007T223915Z-dd5502.jsonl). Fixed: when the forced reply has no usable final, the fallbacks are the answer dict, then an answer variable, then the reply's prose with every code block and<think>block removed (dropped if it reads like a plan or only introduces code), then an empty answer. A code block is never returned. The same logic covers a root timeout, where the call is bounded to one attempt with thinking off and a 60 s request timeout, andprotocol="tools", where afinal_answercall orFINALtext is accepted.
Defects in the tasks. The generators were left unchanged, so these rows are scored as the generators score them. The "answerable" columns above leave out the † cells.
- Multivac's key never matches its question. (Fixed after this run: the target is now picked once; see
evals/independent/fixed/round2/.)get_question()andget_answer()each callrng.choiceseparately, so the key describes a different requirement from the one asked about. It mismatched in all 18 seed × size combinations checked. The bug is in the original, in the author's fix and in the MiniMax fix.check.pyonly testsscore(truth) == 1, so it passed. In the rows we read, both modes answered about the requirement that was asked for. - TheDixieFlatline's answer is sometimes absent from the context. (Fixed after this run by its author: starting offices are now stated.) The starting offices are never written into the context. If the final holder never moved or confirmed an office, the answer cannot be found. That happened in 3 of the 5 cells run: seed 0 small and medium, and seed 1 small.
- SELMA's "Unassigned" key can contradict the text. (Fixed after this run: unassignment now emits an email.) A random
unassignevent produces no email, so in 3 of the 5 cells run (seed 0 large, seed 1 small and medium) the key is "Unassigned" while the latest dated email names a lead. Both modes named that lead. - MasterControl's scorer rejects a name written inside a sentence. (Fixed after this run, along with unformatted amounts like
39195.) Its name regex runs withIGNORECASE, so"Priya Patel was the only employee flagged…"matches as one long "name", and a correct answer gets 0.5. That is MasterControl small seed 0 for rlm. The cell is counted as answerable and the 0.5 stands.
Verdict. On tasks written by other models, the harness does not match the
self-authored results. Below the window, it is worse than just prompting the model.
Above the window, it is the only way to get an answer at all, but it gets one on only
about half of the answerable medium cells and a fifth of the large ones. Its answers
are reliable on tasks that a few greps can solve, and weak on multi-document
reconciliation. The clear levers are the turn budget and getting Qwen to delegate
reading to llm_query, plus the two forced-finish bugs (since fixed). Any change to those must be
measured on a fresh seed or new tasks, not on these rows.
Frozen harness: commit 458e2f76c29ce668f898cc39d15b0ae6a8eb087c. We re-ran the same
grid once the defects found above were fixed. Two things changed since the first run:
- The harness: both forced-finish bugs were fixed (PR
#34). The forced finish now
always asks the model once, shows it the REPL state, and never returns code. Batch
metadata records
src/reclamo/{client,parsing,prompts,rlm}.pyas changed frome17580f. Nothing else changed. - The tasks: four defects were fixed (PR #35): the Multivac key, the missing starting offices in TheDixieFlatline, SELMA's unassignment email and MasterControl's scorer. GLaDOS, SHODAN, Cerebex and Neuromancer are unchanged, and for them every seed × size produced the same context and key as in the first run.
Everything else was the same as the first run: the default pluto profile, the fence
protocol, depth 1, 20 turns, max_timeout=900 s and run_eval.py unchanged. Nothing was
tuned. The grid was also the same: 8 tasks × small/medium/large at seed 0, plus seed 1 for
small and medium, for 80 rows. Batches ran one task × size at a time under the pluto lock,
all small, then medium, then large. The run took 3 h 38 min on 2026-10-07/08. Raw rows are
in runs/independent-rerun-20261007.json, with one trajectory per rlm row in runs/
(not in git).
Each cell shows seed 0 / seed 1. The first run's † cells were tasks whose key could not
be reached. The fixes removed them for Multivac and TheDixieFlatline. ‡ marks SELMA
cells that are still ambiguous because of a remaining SELMA defect (see "Defects in the
round-2 tasks" below). The scores are what each generator's score() returned.
| Task (author model) | Size | rlm score | plain | rlm turns | rlm sub-calls | rlm s | rlm stop |
|---|---|---|---|---|---|---|---|
| GLaDOS (grok-4.6) | small | 1 / 0.80 | 1 / 1 | 20 / 20 | 0 / 0 | 541 / 330 | answer_dict / answer_dict |
| medium | 0.60 / 0.60 | does not fit (~148K tokens) | 20 / 20 | 0 / 0 | 215 / 316 | out of turns / out of turns | |
| large | 1 | does not fit (~606K tokens) | 20 | 0 | 273 | out of turns | |
| SHODAN (gpt-6-astra) | small | 0 / 0 (error) | 1 / 0 | 20 / – | 1 / – | 183 / 251 | out of turns / error |
| medium | 0 / 0 | does not fit (~89K tokens) | 20 / 20 | 0 / 22 | 216 / 995 | out of turns / out of turns | |
| large | 0 | does not fit (~354K tokens) | 20 | 0 | 207 | out of turns | |
| TheDixieFlatline (gemini-3.1-pro-high) | small | 0 / 0 | 1 / 1 | 11 / 14 | 0 / 0 | 167 / 267 | answer_dict / answer_dict |
| medium | 0 / 1 | does not fit (~99K tokens) | 11 / 18 | 0 / 0 | 155 / 257 | answer_dict / answer_dict | |
| large | 0 | does not fit (~400K tokens) | 13 | 0 | 216 | answer_dict | |
| Cerebex (glm-5.3-flash) | small | 0 / 0 | 0.50 / 1 | 20 / 20 | 0 / 1 | 306 / 465 | final / out of turns |
| medium | 0 / 1 | does not fit (~93K tokens) | 19 / 20 | 0 / 0 | 510 / 386 | answer_dict / final | |
| large | 0 | does not fit (~373K tokens) | 20 | 4 | 345 | answer_dict | |
| Neuromancer (deepseek-v4-flash) | small | 1 / 0 | 1 / 1 | 7 / 11 | 0 / 1 | 68 / 438 | final / answer_dict |
| medium | 0 / 1 | does not fit (~87K tokens) | 20 / 14 | 1 / 1 | 604 / 317 | answer_dict / answer_dict | |
| large | 1 | does not fit (~351K tokens) | 20 | 0 | 366 | out of turns | |
| SELMA (nemotron-3-super-120b-a12b) | small | 1 / 0‡ | 0 / 1‡ | 6 / 5 | 0 / 0 | 45 / 39 | answer_dict / final |
| medium | 1 / 0‡ | does not fit (~90K tokens) | 10 / 9 | 0 / 0 | 87 / 65 | final / answer_dict | |
| large | 0‡ | does not fit (~358K tokens) | 7 | 0 | 57 | final | |
| MasterControl (mistral-medium-3.1) | small | 1 / 1 | 1 / 1 | 8 / 10 | 0 / 0 | 78 / 94 | answer_dict / answer_dict |
| medium | 1 / 1 | does not fit (~87K tokens) | 11 / 10 | 0 / 0 | 176 / 186 | answer_dict / answer_dict | |
| large | 1 | does not fit (~344K tokens) | 9 | 0 | 166 | answer_dict | |
| Multivac (hermes-4-405b) | small | 1 / 0.50 | 1 / 0.50 | 5 / 6 | 0 / 0 | 38 / 32 | final / final |
| medium | 1 / 1 | does not fit (~106K tokens) | 7 / 7 | 0 / 0 | 61 / 52 | answer_dict / answer_dict | |
| large | 1 | does not fit (~429K tokens) | 8 | 0 | 68 | answer_dict |
Accuracy by size in the re-run. A plain prompt that does not fit scores 0. The error row scores 0.
| Size | Cells | rlm mean | rlm exact | plain mean | plain exact | plain does not fit | Cells without ‡ | rlm mean | rlm exact | plain mean | plain exact |
|---|---|---|---|---|---|---|---|---|---|---|---|
| small (~60K) | 16 | 0.46 | 6/16 | 0.81 | 12/16 | 0/16 | 15 | 0.49 | 6/15 | 0.80 | 11/15 |
| medium (~300K) | 16 | 0.57 | 8/16 | 0.00 | 0/16 | 16/16 | 15 | 0.61 | 8/15 | 0.00 | 0/15 |
| large (~1.2M) | 8 | 0.50 | 4/8 | 0.00 | 0/8 | 8/8 | 7 | 0.57 | 4/7 | 0.00 | 0/7 |
Comparison with the first run, same cells (mean / exact):
| Size | rlm, first run | rlm, re-run | plain, first run | plain, re-run |
|---|---|---|---|---|
| small (~60K) | 0.38 / 5 of 16 | 0.46 / 6 of 16 | 0.59 / 9 of 16 | 0.81 / 12 of 16 |
| medium (~300K) | 0.36 / 5 of 16 | 0.57 / 8 of 16 | 0.00 (does not fit) | 0.00 (does not fit) |
| large (~1.2M) | 0.15 / 1 of 8 | 0.50 / 4 of 8 | 0.00 (does not fit) | 0.00 (does not fit) |
Split by whether the task changed (rlm over all three sizes; plain at small, the only size where it fits):
| Tasks | rlm, first run | rlm, re-run | plain small, first run | plain small, re-run |
|---|---|---|---|---|
| Unchanged (GLaDOS, SHODAN, Cerebex, Neuromancer) | 0.33 / 5 of 20 | 0.40 / 6 of 20 | 0.81 / 6 of 8 | 0.81 / 6 of 8 |
| Fixed (TheDixieFlatline, SELMA, MasterControl, Multivac) | 0.33 / 6 of 20 | 0.62 / 12 of 20 | 0.38 / 3 of 8 | 0.81 / 6 of 8 |
Findings:
-
Most of the gain comes from the task fixes, not the harness. The rlm improvement is concentrated in the four fixed tasks, where exact answers went from 6 of 20 to 12 of 20. Multivac went from 0 to exact on 4 of 5 cells, and its fifth cell is right but under-scored (see below). MasterControl is now exact everywhere. On the four unchanged tasks, rlm went from 5 to 6 exact out of 20, and plain scored exactly as before on all 8 small cells. That rlm change is two exact cells gained (GLaDOS large and Neuromancer large) and one lost (Neuromancer small seed 1). The contexts were identical, so this is consistent with run-to-run sampling noise.
-
The forced-finish fixes did not rescue any answer. There were 9 forced finishes (
forced_finishevents), all from running out of turns, against 15 in the first run. Of these, 8 hadsource=model: the forced reply contained a validFINAL(...). The other wassource=reply_text. None usedanswer_dictorvariable, andshownwas empty in all 9, so no answer dict or answer variable existed to show. Two of the 9 were exact, the same count as the first run's 15:- GLaDOS large (
runs/20261008T071548Z-de1caa.jsonl): the reply was a modelFINAL. - Neuromancer large (
runs/20261008T073313Z-48c9be.jsonl): the reply wasFINAL ≈ **$91,049.09**, which does not parse asFINAL(...). The new fallback kept its prose, the same total the first run lost to a staleresultvariable.
This time, however, that trajectory made no
llm_querycall and set noresult, and the reply contained no code. The old code would have returned the same text. Bug 1 (a stale variable) never had a variable to act on, and bug 2 (code returned as the answer) never had code to strip. The fixes are therefore untested by this run rather than shown to help. No answer contained a code block. - GLaDOS large (
-
Below the window, plain still wins clearly. At ~60K characters, plain was exact on 12 of 16 cells and rlm on 6. That 12 is best-of-2 on the answer key; as a single attempt plain was exact on 9 (thinking on) or 8 (thinking off), still ahead of rlm. Plain took a median 69 s for the chosen call (167 s for both variants), against 175 s for rlm. TheDixieFlatline is now answerable, and it shows the gap: plain was exact on both seeds, and rlm was wrong on both. On small seed 0, rlm traced the renames correctly. It then took an Auditor's "going to be moved to Room 112 tomorrow" as the final location (
runs/20261008T044627Z-6d2fa3.jsonl). -
Above the window, the harness is still the only option, and it is better than in the first run. It was exact on 8 of 16 medium and 4 of 8 large cells, against 5 of 16 and 1 of 8 in the first run. It is reliable on MasterControl and Multivac (exact on every medium and large cell) and still never solves SHODAN (0 of 5). GLaDOS medium keeps getting 0.60: a wrong amount, or a right amount with no certifying member. Cerebex large was wrong despite 4 sub-calls.
-
The turn cap is still the main harness-side limit. 15 of 40 rlm rows used all 20 turns, and 9 of them ended in a forced finish (2 exact). SHODAN medium seed 1 made 22
llm_querycalls, then gave up with all-zero credits. Its forced reply says "the aggregation step never ran" (runs/20261008T061131Z-c01315.jsonl). -
Qwen still rarely delegates. 7 of 40 rlm rows made any sub-call, against 5 in the first run, and one row accounts for 22 of the 31 calls.
-
Protocol is still clean, with one aborted row. 536 code executions produced 1 syntax error and 8 execution errors, with 0 rejected finals and 7 protocol slips. One row, SHODAN small seed 1, is an error row (
runs/20261008T043757Z-64a70a.jsonl): turns 13–15 raised aSyntaxErrorand then twoIndexErrors, andmax_errors=3raisedRLMErrorLimit. That exception ends the run with no forced finish, so 15 turns of REPL state produced no answer. This is a new harness finding. It is recorded here and not fixed. -
max_timeoutcan overrun by a turn plus the forced finish. SHODAN medium seed 1 took 995 s against the 900 s cap. The deadline is checked when each turn starts. Turn 20 started at 852 s and ended at 929 s, and the forced finish after running out of turns is not time-bounded (only the timeout path is). No run looped or stalled, and none had to be killed.
Defects in the round-2 tasks. The scores above stand as the round-2 generators compute
them. The round-3 audit found both of these independently, and they were fixed on main
in PR #46, which merged during
this run's large phase. This run used the round-2 versions pinned at 458e2f7 throughout:
run_eval.py checks each generator's sha256 against that manifest.
- SELMA's round-2 fix is incomplete (‡). The key follows the emails' order in the
document, but each email's displayed
Date:is a random day within the event's year (_rand_date(rng, year, year)). The latest-dated email can therefore contradict the key. In all three "Unassigned" cells that were run, the latest-dated AI Ethics email assigns a lead. These are small seed 1, medium seed 1 and large seed 0, the same three cells as the first run's †. In large seed 0, for example, "Position Now Unassigned" is dated 2032-01-20 and an assignment to Taylor Jackson is dated 2032-04-13. rlm sorted by date and answered Taylor Jackson (runs/20261008T073921Z-ab1184.jsonl). Over seeds 0–9, 12 of 30 seed × size cells have this conflict. The round-2 self-test passed because it finds the last email by position, not by date. - Three scorers reject correct answers written in markdown or as phrases. This is the
same kind of defect that MasterControl had:
- SELMA splits the answer on whitespace, so
**Skyler Martin**keeps its asterisks and scores 0. On SELMA small seed 0, plain's thinking variant gave that correct answer in bold. Its no-thinking variant said "Unassigned", so the cell scored 0. - Multivac's status regex rejects both "status is Blocked" and "Status: Blocked". On Multivac small seed 1, both modes gave the correct status and full description and scored 0.50.
- GLaDOS's outcome check rejects "DID pass". On GLaDOS small seed 1, rlm had the right amount and name and scored 0.80.
- SELMA splits the answer on whitespace, so
Verdict. After the fixes, the picture is the same in kind and better in degree.
Below the window, the plain model is still clearly better: 12 exact against 6. Above it,
the harness answers about half the cells exactly, where plain answers none: 8 of 16
medium and 4 of 8 large, up from 5 and 1. The improvement comes from repairing the
tasks; on unchanged tasks it is within noise. The forced-finish fixes were not
exercised. The open levers are the same as before: the 20-turn budget, delegation to
llm_query, and multi-document reconciliation (SHODAN). Two new harness questions came
up: whether max_errors should end in a forced finish rather than an exception, and
whether the forced finish should be time-bounded. Both need measuring on new seeds, not
on these rows. These numbers are also for the round-2 task set. The round-3 set in PR #46
changes the keys and scorers of all eight tasks, so it needs its own run.
These tasks, like the benchmark, protocol and depth-2 tasks below, were written by the harness's own authors. For tasks written by other models, see Independent evaluation.
examples/eval.py --profile pluto, seed 0, run 2026-10-07 against
huihui-qwen3.8-flash-next-abliterated (Qwen3.8-Flash-Next, UD-Q4_K_XL) served by
Strata on pluto (V100 32 GB + P100 16 GB + RTX 3060 12 GB, one request at a time,
32K resident KV). Default pluto profile: root thinking on, sub-calls thinking off,
subcall_chars=12000, depth 1.
| Task | Result | Turns | Sub-calls | Tokens | Seconds |
|---|---|---|---|---|---|
| needle (1,000,000 lines, ~30 MB) | found | 4 | 0 | 7,022 | 18.6 |
| oolong_lite (300 tickets, 5 categories) | exact: total abs error 0 | 6 | 3 | 21,279 | 127.7 |
Notes:
- The needle is solvable with code alone; Qwen scans for the odd line and never needs a sub-call. The context is ~1,000x the model's resident window.
- OOLONG-lite cannot be grepped. Qwen batched the tickets into 3
llm_querycalls of ~100 tickets each and aggregated in the REPL, instead of one call per ticket. That is the over-calling failure the RLM paper reports for Qwen3-Coder (hundreds of calls per task), held off here by the batching instruction plus hard caps in code. - One run per task; treat these as smoke results, not a benchmark.
The tasks in this benchmark were written by the harness's own authors. On tasks written by other models the harness does much worse; see Independent evaluation.
#19. examples/bench.py --profile pluto --seeds 3, run 2026-10-07 against the same server and default profile
as above. Raw rows: runs/bench-20261007T094811Z.json (not in git).
What each mode means:
- rlm is the harness.
- plain is one chat call with the whole context in the prompt. Since #57 it is one
attempt with thinking as the root role sets it;
plain-thinkandplain-nothinkforce thinking on or off, each as its own row. - does not fit means the plain prompt would not fit pluto's usable window, so the plain model cannot attempt the task. The usable window is 28,672 tokens: 32K resident KV minus an output reserve.
The token estimate counts one token per digit and 3.5 characters per token otherwise. The plain rows in this table predate #57: each was the better of two variants, thinking on with a 16K output cap and thinking off, chosen against the answer key. That is an oracle the harness did not get, so these plain results are an upper bound (see the correction at the top of Results). The 16K output cap stays; the selection is gone.
| Task | Size | Mode | Result over 3 seeds | Median s | Median sub-calls |
|---|---|---|---|---|---|
| needle | 2,000 lines | rlm | 3/3 found | 13.7 | 0 |
| needle | 2,000 lines | plain | 3/3 found | 0.6 | – |
| needle | 10,000 lines | rlm | 3/3 found | 13.2 | 0 |
| needle | 10,000 lines | plain | does not fit (~74K tokens) | – | – |
| needle | 100,000 lines | rlm | 3/3 found | 16.3 | 0 |
| needle | 100,000 lines | plain | does not fit (~1.5M tokens) | – | – |
| needle | 1,000,000 lines | rlm | 3/3 found | 15.9 | 0 |
| needle | 1,000,000 lines | plain | does not fit (~16M tokens) | – | – |
| oolong_lite | 100 tickets | rlm | exact 3/3 | 84 | 4 |
| oolong_lite | 100 tickets | plain | exact 3/3 | 85 | – |
| oolong_lite | 300 tickets | rlm | exact 3/3 | 117 | 3 |
| oolong_lite | 300 tickets | plain | exact 3/3 | 292 | – |
| oolong_lite | 1,000 tickets | rlm | exact 2/3; one seed off by 34 | 380 | 16 |
| oolong_lite | 1,000 tickets | plain | 2 of 3 do not fit; the one that fit was wrong (total error 360) | – | – |
| longdoc_qa | 100 sections | rlm | 9/9 questions | 38 | 0 |
| longdoc_qa | 100 sections | plain | 9/9 questions | 82 | – |
| longdoc_qa | 300 sections | rlm | 9/9 questions | 80 | 0 |
| longdoc_qa | 300 sections | plain | does not fit (~53K tokens) | – | – |
Findings:
- Where the context fits, the plain model is just as accurate. At 100 and 300 tickets, the 100-section QA and the 2,000-line needle, plain was exact every time. The harness is not smarter on small inputs.
- Where it doesn't fit, only the harness can answer. That covers the needle from 10K to 1M lines, 1,000 tickets and 300 sections. The harness found the needle all 12 times at up to ~16M estimated tokens, in about 16 s each. It answered all 18 QA questions and got 1,000 tickets exact on 2 of 3 seeds.
- The harness is often faster even when plain fits. It took 117 s against 292 s at 300 tickets, and 38 s against 82 s on the 100-section QA. It keeps the root prompt small instead of thinking over the whole document in one call. The exception is the tiny needle, where plain answers in under a second.
- The first plain baseline was unfair. Two of its three 100-ticket failures were cut-offs: thinking ran into the output limit. A best-of-two rule replaced it, and plain went to 3/3, but that rule picked the variant using the answer key and is itself unfair in plain's favour. Since #57 the cut-off is handled by the generous output cap alone, with one attempt.
- Sub-calls stayed small. 0 to 16 per run. 16 was at 1,000 tickets, still about 60 tickets per call rather than one per ticket.
- The one miss was oolong_lite at 1,000 tickets, seed 2, total error 34. The other two seeds were exact. This is the size where per-batch classification errors start to add up.
- Caveats. These are synthetic tasks written by the harness's authors. See
#22 and
evals/independent/for the tests written by others. Aplain_truncatedmode (the head of the context, cut to fit) exists inbench.pybut is not reported: an early version under-estimated digit-heavy text and sent prompts too long for the 300 s request timeout. It is fixed, but has not been re-run.
#20. --protocol tools
(config protocol = "tools") has the root model act through OpenAI-style function
calls instead of fenced code and FINAL(...) text. execute_python(code) returns the
REPL output, with the same truncation and sub-call caps. final_answer(answer=... | variable=...) finishes; it takes exactly one of the two, and variable must exist.
The history stays append-only. Each turn adds the assistant message with its
tool_calls, then one tool message per call in call order, then a user message with
any notes and the next Turn i/N line. The guards from the fence protocol still apply:
a final sent next to code is rejected, a plan-like final is rejected once, the
re-verify and decompose nudges fire, and the forced finish and the timeouts work the
same way.
A reply with no tool call gets a nudge. A fenced block or a FINAL(...) written as text
is honoured once and logged as a protocol slip. The fence prompt is unchanged. Strata
returns structured tool_calls for Qwen3.8-Flash-Next, so there is no text-parsing
fallback.
Text-form tool calls (#48).
Some servers do not parse a model's native tool-call syntax and leave it in the message
text. Poolside Laguna S 2.1 does this: it writes <tool_call>repl\n<code></arg_value></tool_call>
in a fence-protocol run. The fence parser accepts that form as a code block (names
repl, python, execute_python). The code ends at the first </arg_value>,
</value>, </repl>, </tool_call> or </think>, or at the end of the reply, and
trailing arguments such as description are ignored. If a reply's first code is a
<tool_call>, only that block runs and the rest of the reply is discarded. Laguna
sometimes follows the block with REPL output it made up, often in a ```repl
fence, and then more calls and a FINAL built on that output. A reply that opens with
a fence is handled as before, with any later tool calls run in document order.
<tool_call>FINAL(...) counts as FINAL(...), under the same rules: next to code that
has not run yet, it is rejected. Strict parsing (```repl fences only, the
upstream protocol) ignores this form.
Run 2026-10-07 on the same server and profile as above. Both protocols were
interleaved per seed, and each row took the pluto lock on its own. bench.py --modes rlm --protocol {fence,tools} --seeds 3 gives the raw rows in
runs/proto-bench-20261007.json (not in git). "Syntax errors" counts executed blocks
whose error was a SyntaxError. "Slips" counts the protocol slips above, plus fence
replies with neither code nor a final.
| Task | Size | Protocol | Result over 3 seeds | Final rejections | Slips | Syntax errors / execs | Median turns | Median s |
|---|---|---|---|---|---|---|---|---|
| needle | 100,000 lines | fence | 3/3 found | 0 | 0 | 0/6 | 3 | 14.9 |
| needle | 100,000 lines | tools | 3/3 found | 0 | 0 | 0/6 | 3 | 18.1 |
| oolong_lite | 300 tickets | fence | exact 3/3 | 0 | 0 | 0/17 | 7 | 121 |
| oolong_lite | 300 tickets | tools | exact 3/3 | 0 | 0 | 0/21 | 7 | 244 |
| oolong_lite | 1,000 tickets | fence | exact 3/3 | 0 | 0 | 0/27 | 10 | 454 |
| oolong_lite | 1,000 tickets | tools | exact 3/3 | 0 | 0 | 0/14 | 5 | 304 |
| longdoc_qa | 300 sections | fence | 6/9 questions (one seed out of turns) | 0 | 0 | 0/44 | 15 | 134 |
| longdoc_qa | 300 sections | tools | 9/9 questions | 0 | 0 | 0/30 | 11 | 94 |
The independent eval set (evals/independent/fixed/active/, size medium, seed 0) was
run as a protocol comparison only; nothing was tuned on it, because it is reserved for
the frozen #22 eval. Raw
rows are in runs/proto-indep-20261007.json, with max_timeout=900.
| Protocol | Mean score (8 tasks) | Exact | Final rejections | Slips | Syntax errors / execs | Median turns | Median s |
|---|---|---|---|---|---|---|---|
| fence | 0.28 | 2/8 (MasterControl, SELMA) | 0 | 2 (no code, no final) | 0/113 | 19.5 | 231 |
| tools | 0.40 | 3/8 (MasterControl, SELMA, Neuromancer) | 0 | 0 | 0/125 | 19 | 209 |
Findings:
- The protocol errors #20 set out to remove do not occur any more. Over 24 bench runs and 16 independent runs, neither protocol produced a syntax error or a rejected final. Qwen never slipped out of the tools protocol: no text-only replies, no fenced code, no malformed arguments. The fence protocol had two turns with neither code nor a final, both on independent tasks the model failed anyway. The prompt and parser tuning in #7 already removed what the tools protocol would have fixed.
- Accuracy is the same within noise. Each protocol's extra wins come from a single run. Fence lost one longdoc_qa seed by searching for 20 turns, after which the forced finish returned its reasoning text. That task got 9/9 with fence in the #19 benchmark. Tools won Neuromancer only through the forced finish, which picked up an answer already sitting in a REPL variable. Both protocols failed the same five independent tasks; three of them ran out of turns.
- Speed is mixed. Tools was faster on longdoc_qa (94 vs 134 s), on 1,000 tickets (304 vs 454 s) and on the independent set (209 vs 231 s). It was twice as slow on 300 tickets (244 vs 121 s): it chose 50-ticket batches and a verification pass where fence used three 100-ticket calls. That is a strategy difference, not a protocol error. Tool turns also cost more prompt tokens, because of the tool schemas and the JSON-escaped code.
- The default stays
fence. Tools is not clearly better: it has no error rate to improve and it ties on accuracy.--protocol toolsis supported and tested for servers or models whose fenced-code output is less reliable. Its stats show up in every trajectory'sfinalrecord and in bench rows, so the comparison can be re-run cheaply. With three seeds and one run per independent task, these are observations, not significance tests.
#18. Same server and
profile, seed 0, 2026-10-07, examples/nested.py (12 departments, 308,990 chars).
At depth 2 the system prompt gains the rlm_query(question, data) tool and a
"delegate" example, and a sub-call may start a nested RLM with data as its own
context. Sub-calls include the children's own llm_query calls; child turns are the
root turns taken inside nested RLMs. Trajectories: runs/20261007T094618Z-6d710f,
100122Z-16afdc, 103540Z-49d014, 104924Z-80a218, 110425Z-3b9b20 (.jsonl).
| Task | Depth | Result | Turns | Sub-calls (child turns) | Tokens | Seconds |
|---|---|---|---|---|---|---|
| nested | 1 | exact: 12/12 counts, "most" right | 20 (forced finish from the answer dict) | 5 (0) | 167,301 | 275.5 |
| nested | 2, v0.1 rlm_query(prompt) |
abs error 43, "most" right; never called rlm_query |
19 | 3 (0) | 180,529 | 426.0 |
| nested | 2, rlm_query(question, data) |
abs error 73, "most" wrong; never called rlm_query |
16 | 4 (0) | 140,675 | 320.3 |
| nested | 2, --delegate |
11/12 departments exact, then max_timeout=900 hit during the 12th; no answer returned (fixed below) |
4 (+84 child) | 16 (84) | 328,933 | 900.5 |
| oolong_lite (300 tickets) | 2 | exact: total abs error 0; never called rlm_query |
7 | 6 (0) | 24,910 | 137.6 |
Findings:
- Qwen does not choose to recurse. In three depth-2 runs it never called
rlm_query, even after the tool was redesigned to take(question, data)and the prompt gained a delegate example. It did at depth 2 what it does at depth 1: dedupe the 477 complaints into ~110 templates, classify those with a few batchedllm_querycalls, and replay the open/close/reopen events in code. The two depth-2 misses were model slips, not recursion: one 161-line batch came back with 186 labels (misaligned counts), and in the other run a spot-check turn decremented the result dict to zero and the next turn submitted it, the paper's E.2 failure, live. - When told to delegate, nested RLMs work but are slow. 11 of the 11 children that
finished were exactly right (4 to 11 turns each, 37 to 128 s, format discovery and the
event replay in code, the semantic judgement inline in a thinking turn). At ~80 s per
department the 12 children need ~16 min against 4.6 min for depth 1 on a server that
takes one request at a time, and the run hit the 15 min cap on the 12th. The harness
then threw away the eleven results: the child's timeout reached the root as a
RuntimeErrorinside its loop and the root's own deadline check raised with no partial answer. Fixed in the same PR: a root timeout now gets the forced finish that running out of turns already had (REPL value first, else one "out of time" call). - Depth 1 is not free either: 20 turns, mostly format discovery across 12 logs, and it
ran out of turns with the answer already in the dict because
FINAL_VAR(answer['content'])was rejected as "no such variable" (also fixed). - The default stays
max_depth=1. On pluto recursion costs about 3x the wall clock and 2x the tokens for the same answer, and the model will not use it unless the task tells it to. The per-department accuracy is the encouraging part; revisit with a server that runs children in parallel or a smaller per-child turn budget. One seed per row (pluto was shared, with 6 to 24 min lock waits per run), so these are observations, not a benchmark.
uv sync
uv run ruff check
uv run ruff format --check
uv run pytest # live tests (-m live) are skipped by default- Alex L. Zhang, Tim Kraska and Omar Khattab, Recursive Language Models, arXiv 2512.24601.
- alexzhang13/rlm and alexzhang13/rlm-minimal, both MIT. This project follows their design; see NOTICE and THIRD_PARTY_LICENSES.md.
Apache-2.0. See LICENSE.