Repository navigation
LIMITS-001: runtime throughput is the bootstrapping blocker — ~650-850K instr/sec, no JIT #173
Description
Activity
- addedseverity:lowMinor surpriseMinor surprise
on Jun 7, 2026 - added a commit that references this issue
on Jun 7, 2026 - addedcycle:v5.0v5.0 cyclev5.0 cycletier:4-deferred-to-v5Tier 4: deferred to v5.0Tier 4: deferred to v5.0
on Jun 8, 2026 Triage 2026-08-17 (against v5.0.2). The measurement in this issue is wrong — the ceiling is roughly 2.3x higher
than stated. The conclusion still holds; the number should not be quoted.Measured on 5.0.2, a tight arithmetic loop through the real embedding API:
5,100,010 instructions in 11.24 s -> ~454,000 instr/secagainst the ~200K instr/sec this issue claims. Measured on a machine that is
currently slower than normal (see the timing note inCLAUDE.md), so the true
figure is likely higher still.Everything structural in the issue remains true: the VM is a CPython dispatch loop,
there is no JIT, no loop-invariant hoisting, no escape analysis. Only the throughput
figure needs correcting, and it is the kind of number that gets quoted in
positioning documents, which is why it is worth fixing rather than leaving.Refiling this as the bootstrapping blocker
Re-measured on 5.0.4, and reframing what this issue is. It has been carried as a
generic performance note atseverity:low; it is in fact the thing standing
between Nodus and self-hosting, whichdocs/language/LANGUAGE_VISION.mdnow
states as a decided direction rather than an aspiration.Current measurement
Executed rather than estimated, embedded runtime with no step or time limit:
SRC = ''' fn main() { let i = 0i let acc = 0i while (i < 1000000i) { acc = acc + i; i = i + 1i } print("acc=\(acc)") } ''' rt = NodusRuntime(timeout_ms=None, max_steps=None)
acc=499999500000 1,000,000 loop iterations in 54.13s stats: {'instructions_executed': 17000021, 'coroutines_spawned': 0}≈ 314,000 instructions/sec — about 17 instructions per trivial loop
iteration. Somewhat better than this issue's original ~200K estimate, and the
same order.Why this is the bootstrapping blocker specifically
The three other prerequisites the roadmap names are in good shape:
Prerequisite State Stable language semantics core surfaces are Stable; 18 adversarial corpus reads produced zero findings against lexer/parser/VM/typesStable bytecode instruction set 49 opcodes, BYTECODE_VERSION4, frozen since v1.0 and gate-enforcedReliable module system import/exportstable; #158 (no privacy) is an annoyance, not a barrierSufficiently expressive stdlib the weak one — no string slice, no byte type (#170), no closure upvalue mutation (#156) And the shape of the task is already demonstrated:
examples/expr_compiler.nd
is a character-level lexer, recursive-descent parser and tree-walking evaluator
written entirely in Nodus, 231 lines.So the gap is not can the language express a compiler — it demonstrably can.
The gap is arithmetic. The pipeline that would need rewriting is ~2,300 lines of
lexer/parser/AST plus ~1,900 lines of compiler. A Nodus lexer without a slice
operator indexes character by character; at ~314K instr/sec, compiling a single
mid-sized source file lands in the seconds, and a self-hosted compiler compiling
itself is the workload that has to be affordable.Nothing in the other three prerequisites changes that. Only this does.
On the fix directions
The three listed still look right, and the middle one is better positioned than
when this was filed:- PyPy backend — no code changes, 5–10× on dispatch-heavy loops. Cheapest
thing to actually measure, and worth measuring before choosing anything else:
if it lands near 3M instr/sec the calculus changes completely. - Native VM loop — PyO3 scaffolding exists and is now proven in production,
not just present:nodus-native-memory-enginev0.1.1 ships a Rust extension
with a pure-Python fallback. The pattern of "native where it pays, Python
fallback everywhere" is established in this ecosystem rather than theoretical. - Selective compilation — still the largest IR investment, still last.
Suggested first step is (1) as a measurement rather than a commitment: run the
same 1M-iteration probe under PyPy and record the number. It is an afternoon,
and it decides whether (2) is necessary or merely nice.What changes on this issue
- severity:low → severity:medium. Not high: nothing is broken, no user is
blocked today, and the workarounds listed are real. But "low" reads as this
does not matter, and it is the single blocker on a stated direction of the
language. - Adding
type:design-question, because choosing among the three paths is a
decision, not an implementation task. - The "None of these are planned for v4.x" line is now just true-and-stale; v4.x
is three majors back.
Provenance
Surfaced while auditing how visible the bootstrapping goal is and how close the
language actually is to it — the answer being that it was documented nowhere a
reader would find it, and that this issue is what stands in the way. Docs fixed
separately; this is the tracking half.- PyPy backend — no code changes, 5–10× on dispatch-heavy loops. Cheapest
- addedseverity:mediumUX broken but workaround existsUX broken but workaround existstype:design-questionLanguage-level design decision, not a simple fixLanguage-level design decision, not a simple fixand removedseverity:lowMinor surpriseMinor surprise
on Aug 19, 2026 - changed the title
[-]LIMITS-001: Python interpreter throughput ceiling — ~200K instr/sec, no JIT[/-][+]LIMITS-001: runtime throughput is the bootstrapping blocker — ~314K instr/sec, no JIT[/+]on Aug 19, 2026 Measured: PyPy is ~23× faster, and Nodus runs on it unmodified
This issue's first fix direction was "PyPy backend — no code changes, 5–10× on
dispatch-heavy loops", and my earlier comment argued it should be measured
before committing to a native VM, because if it landed near 3M instr/sec the
calculus changes completely. It lands well past that.The numbers
Same 1M-iteration probe (17,000,021 instructions), same machine, three trials
each, best and median reported because a single sample on this box is not
trustworthy — see the correction below.runtime best median vs CPython 3.11 CPython 3.11.9 407,097 401,694 — CPython 3.14.2 454,600 383,872 ~1× (noise) PyPy 7.3.23 (3.11.15) 10,048,618 9,384,307 ~23× PyPy trial 1 was 3.06s against 1.69s and 1.81s for trials 2 and 3 — JIT warmup,
so the warm figure is the honest one for a long-running compile.Verified as real work rather than skipped work: PyPy produces
acc=499999500000and executes the same 17,000,021 instructions.Correcting my own figure
My earlier comment reported ~314K instr/sec from a single run, and that number
went intoLANGUAGE_VISION.mdand the README. Re-measuring gives 365K–454K on
CPython across six trials. The 314K was a slow sample on a noisy box, not a
wrong method — but it was one sample presented as a measurement, which is the
thing I should not have done. The docs should say ~400K, and I will correct
them.It does not change any conclusion. The ratio is what matters and it is ~23× either
way.CPython 3.14 buys nothing
Worth recording so nobody repeats it: 3.14 is within noise of 3.11 on this
workload. Upgrading the interpreter is not a free win here.Nodus needs no changes to run on PyPy
The only dependency is
tzdata, which is pure data.pip install tzdata, set
PYTHONPATH, andrun_sourceworks first try.One thing blocks it, and it is a real bug either way
A 65-test subset (task graph, workflow DSL, state policies, coroutines, channels)
gives 9 failed, 65 passed — and all nine share a single root cause:cannot commit transaction - SQL statements in progressSQLiteWorkflowStorenever closes the cursorsconn.execute(...)returns. On
CPython refcounting frees them immediately; PyPy's GC does not, so the statement
is still open at commit. Filed as #516.That is not a PyPy incompatibility to work around — it is the store depending on
when CPython happens to free an object for correctness, which is a latent
defect on CPython too. Coroutines, channels, task graph and state policies all
pass unmodified.What this means for the three fix directions
- PyPy backend — now the obvious first move. ~23×, no code changes, one
blocking bug already identified and small. It does not need to become the
supported runtime to be worth having: adding it to CI would catch the whole
class of refcount-dependent resource bugs that CPython structurally cannot. - Native VM loop — should wait. It is the largest investment on the list and
PyPy may make it unnecessary. Reassess after SQLiteWorkflowStore leaks cursors: the store depends on CPython refcounting to commit #516 and a full suite run. - Selective compilation — unchanged, still last.
What this means for bootstrapping
At ~400K instr/sec a self-hosted compiler compiling itself was the blocker this
issue describes. At ~9.4M it is a different conversation — not solved, but no
longer the thing standing in the way. The next honest step is a real workload
rather than a synthetic loop: runexamples/expr_compiler.ndover a large input
under both runtimes and compare. A tight arithmetic loop is the case a JIT is
best at, so 23× is an upper bound, not a promise.Method
# 33 MB, extracted, no install pypy.exe -m pip install tzdata $env:PYTHONPATH="C:/dev/Coding Language/src" pypy.exe bench.py # 3 trials, NodusRuntime(timeout_ms=None, max_steps=None)
The probe is the one from my earlier comment, with the
printremoved so output
handling is not in the measurement and with three trials instead of one.- PyPy backend — now the obvious first move. ~23×, no code changes, one
- added a commit that references this issue
on Aug 20, 2026 2 remaining items
The compiler workload: PyPy is ~4–5×, not ~23×
My last comment measured PyPy at ~23× on the 1M-iteration arithmetic probe and
concluded that the native VM "should wait" because "PyPy may make it unnecessary".
That was measured on the workload a JIT is best at. Re-measured on the workload that
actually matters for bootstrapping, the answer is different enough to change the
conclusion.The workload
examples/expr_compiler.nd— the character-level lexer, recursive-descent parser and
tree-walking evaluator this issue already cites as proof the language can express a
compiler — driven over a non-trivial expression n times:((1 + 2) * (3 + 4)) - -5 / 2 + 3.5 * 2 * (10 - 4 / 2)8,209 VM instructions per iteration. In-process via
run_source, so interpreter
startup is out of the number; three trials per cell, best reported.The numbers
n instructions CPython 3.11.9 PyPy 7.3.23 (3.11.15) ratio 25 205,215 323,764/s 324,095/s 1.00× 100 820,740 337,151/s 1,388,545/s 4.12× 400 3,282,840 319,053/s 1,741,086/s 5.46× 1600 13,131,240 209,252/s 856,436/s 4.09× Same thing through the CLI as whole processes, which folds in startup and is the
number a user would feel:n CPython wall PyPy wall ratio (startup subtracted) 50 5.10 s 3.39 s 1.81× 200 10.17 s 4.25 s 3.14× 800 36.58 s 8.98 s 4.70× Startup itself: CPython 1.88 s, PyPy 1.60 s.
PyPy peaks around 1.7M instr/sec here, against 10.05M on the arithmetic probe —
5.8× lower on the same machine, same interpreter, same day. The JIT needs a hot loop;
a recursive-descent parser is many cold-ish paths, and it warms slowly: at n=25 PyPy
has no advantage at all, and it is still climbing at n=400.Do not read the n=1600 row as a decline. That row is box noise — a later run of
the same n=1600 case on CPython gave 324,640 instr/sec, in line with the plateau,
and the trial spread on that row was 62.75 / 127.15 / 69.01 s. This machine's timing
instability is documented inCLAUDE.mdand it is fully live. The honest reading of
the table is: CPython ~320K instr/sec flat, PyPy ~1.7M once warm, ratio 4–5.5×.A third of CPython's throughput is not dispatch overhead
While chasing that apparent instability I found something that changes the framing of
this issue. Filed as #522.This issue says the ceiling is "bounded by CPython's own dispatch overhead, not by the
Nodus logic itself". A third of it is neither: the VM emits and retains an event
object for every function call, every return, and every 100 instructions, into an
unbounded list that nothing reads unless a host attaches a sink.Same source, same machine, only
RuntimeEventBus.emitstubbed to a no-op:n bus on bus off speedup events retained 400 302,679/s 454,941/s 1.50× 344,035 1600 324,640/s 490,539/s 1.51× 1,376,119 tracemallocat n=400: 131.4 MB live, 74.4 MB of it the event log — 57% of
everything the run allocates, ~23 bytes per instruction executed.So there is a 1.5× available with no new runtime, no JIT and no native code, and
the memory side is arguably the bigger deal for bootstrapping: at 23 bytes per
instruction, a self-hosted compiler executing a billion instructions retains tens of
gigabytes ofvm_callrecords. Throughput you can wait out; that you cannot.What this does to the three fix directions
- PyPy backend — still worth having, but on the honest workload it is ~4–5×,
not 23×, landing near 1.7M instr/sec. My "if it lands near 3M the calculus
changes completely" test is not met by the compiler workload; it was met only by
the arithmetic probe. It remains the cheapest real gain, it needs no code changes,
and adding it to CI catches the refcount-dependent resource bugs CPython cannot
(which is how SQLiteWorkflowStore leaks cursors: the store depends on CPython refcounting to commit #516 was found). But it does not retire the question. - Native VM loop — I retract "should wait", at least as a conclusion drawn
from the PyPy number. 4–5× on a compiler workload does not make a native dispatch
loop unnecessary. It should be reassessed on its own merits, after The VM retains an event per function call and return, unbounded: 1.5x throughput and 74 MB per 3M instructions, with no consumer #522, not
dismissed on the strength of a benchmark that flattered the alternative. - Selective compilation — unchanged, still last.
And a fourth that was not on the list, which should now be first because it is
strictly cheaper than all three: #522 — stop retaining telemetry nobody reads.
1.5× and the memory wall, in ordinary Python, in the existing runtime.Correcting myself, again
Two things I got wrong on this issue and one method note.
- I reported ~314K instr/sec from a single run; re-measurement gave 365K–454K, and
the docs were corrected to ~400K in docs: re-measure runtime throughput, and record the PyPy result #517. The compiler workload measured here sits
at ~320K, i.e. lower than the arithmetic figure — which is expected and worth
keeping distinct in the docs: the headline number is a hot arithmetic loop, and
real programs are slower. - I concluded the native VM should wait, on a 23× that does not survive contact with
the workload this issue is about. Retracted above.
The method note: the first attempt at this benchmark ran for over ten minutes on a
workload the CLI finished in one second, because the harness passed
filename="compiler_bench.nd"torun_source— andrun_sourceexecutes the file
named byfilenameand discards the source string it was given. It was running the
on-disk copy at 3200 iterations instead of the 25-iteration string it was handed.
That is #521, present since v0.4.0, and I would not have found it without this
benchmark being wrong first.Provenance
Requested follow-up to the PyPy measurement above: run the same comparison on
expr_compiler.ndrather than a synthetic loop, on the reasoning that a tight
arithmetic loop is a JIT's best case and therefore an upper bound rather than a
promise. It was.- PyPy backend — still worth having, but on the honest workload it is ~4–5×,
- added a commit that references this issue
on Aug 20, 2026 - changed the title
[-]LIMITS-001: runtime throughput is the bootstrapping blocker — ~314K instr/sec, no JIT[/-][+]LIMITS-001: runtime throughput is the bootstrapping blocker — ~320-400K instr/sec, no JIT[/+]on Aug 20, 2026 #522 is fixed, and it moves this issue's baseline. Measured on
examples/expr_compiler.nd(400 expressions, 1.96M instructions), same machine,
before and after the change:before after throughput 227,401 instr/sec 477,417 instr/sec events retained 206,382 0 live memory 80.2 MB 0.4 MB That is 2.10×, on CPython, with no JIT and no native VM — because a third to
a half of the time was the runtime buildingvm_call/vm_returnevent objects
that nothing read, and the guard now sits before the object is constructed rather
than after.Two consequences for this issue as written:
-
The "~320-400K instr/sec" figure in the title is now the floor, not the
ceiling. A larger run (7.8M instructions) measures 354K/s post-fix on a
loaded box; the 400-expression run measures 477K/s. Re-measure before quoting
a ceiling. -
The attribution needs revising. This issue says the limit is "CPython's own
dispatch overhead, not the Nodus logic itself". Half of it was neither — it was
telemetry with no consumer. Whatever remains should be re-profiled before it is
used to argue for PyPy or a native VM, for the same reason the PyPy figure was
retracted earlier in this issue: a benchmark that flatters the alternative.
The memory side matters more for the bootstrapping argument here. At ~23 bytes per
instruction retained, a self-hosted compiler executing a billion instructions was
accumulating tens of gigabytes ofvm_callrecords. That wall is gone rather than
raised — retention no longer scales with the run at all.-
- removedtier:4-deferred-to-v5Tier 4: deferred to v5.0Tier 4: deferred to v5.0
on Aug 29, 2026 Re-measured at 5.8.0: the number in this issue is stale by 3–5x
Filed against v4.0.0. Eight minor releases later the ceiling has moved, and nothing recorded it — so the figure people plan around is wrong in the direction that makes the problem look worse than it is.
The title and body already disagreed (
~320-400Kvs~200,000), which is its own signal that neither had been re-derived.Measured
CPython 3.11.9, current
main,NodusRuntime(timeout_ms=None, max_steps=None), best-of-N,instructions_executedfromget_execution_stats()divided by wall clock:shape instr/sec tight arithmetic loop (400K iterations) 1,125K tight arithmetic loop (200K iterations) 935K function calls ( add(a, b)in a loop)970K list list_pushthen indexed read950K ~0.9–1.1M instructions/second, consistent across three unrelated program shapes rather than one favourable microbenchmark.
What that changes in this issue's own impact section
- "CLI default: 200ms wall-clock → ~40,000 compute instructions before timeout fires" —
EXECUTION_TIMEOUT_MSis still 200, but 200 ms now buys roughly 190,000 instructions, not 40,000. - "the 10M step limit (embedded default) fires after ~50 seconds of pure compute" — closer to 10 seconds.
What has not changed
The issue's actual claim stands: it is still a CPython dispatch loop, there is still no JIT, and throughput is still the bootstrapping blocker. A ~5x improvement does not change which of the three fix directions apply.
One of them has since been measured and is worth recording here rather than leaving in a session note: PyPy runs Nodus roughly 23x faster unmodified — the "drop-in, no code changes" path this issue lists first. It is not adoptable today because #516 (a cursor leak) blocks the suite under PyPy. CPython 3.14 was also tried and buys nothing.
Staying open: the limitation is real, only the number was wrong. Re-derive it rather than trusting this comment in another eight releases — the script is four lines against
get_execution_stats().- "CLI default: 200ms wall-clock → ~40,000 compute instructions before timeout fires" —
PyPy measured end to end — it does not deliver, and the ~23x figure does not reproduce
#516 (the cursor leak that blocked this) closed 2026-08-20, so I stood PyPy up and measured rather than inheriting the number. PyPy 7.3.23 / Python 3.11.15 (matching the project's CPython 3.11.9), official win64 build,
.venv-pypy,requirements.txtminusmypy(itsast-serializedependency has no PyPy wheel; irrelevant here since CI type-checks on CPython).tzdatahad to be installed separately — it is a declared dependency inpyproject.tomlbut not inrequirements.txt, andtime_module.pyimportsZoneInfo("UTC")at module import, so Nodus does not start without it.Throughput: 1.5–2.9x, and negative for short programs
Same three shapes, both interpreters, back to back in one window (this box's timing is unstable, so the ratio is worth more than the absolutes):
shape CPython PyPy tight arithmetic loop 638K 1,859K 2.9x function calls 883K 1,319K 1.5x list index + append 760K 1,438K 1.9x Then the same loop at four sizes, which is the more useful cut:
iterations CPython PyPy 50,000 782K 375K 0.48x — PyPy is 2x slower 200,000 598K 821K 1.4x 1,000,000 653K 1,462K 2.2x 4,000,000 897K 1,551K 1.7x It crosses over between 50K and 200K iterations and plateaus near 1.5M instr/sec. The "~23x" figure carried in session notes does not reproduce; I cannot speak to its provenance, and it should not be quoted again without a fresh measurement. This issue's own fix direction — "PyPy backend — drop-in for CPython, can 5–10x throughput on dispatch-heavy loops; no code changes needed" — is not supported by measurement either.
The CLI case is worse, and it is not interpreter startup
CPython PyPy bare interpreter ( -c pass)184 ms 104 ms nodus run hello.nd1,891 ms 4,125 ms PyPy's own startup is faster. The 2.2x penalty is importing and compiling Nodus's 77 modules — code that runs once and never warms up. For
nodus run, whose default deadline is 200 ms, PyPy is strictly worse.The suite is not green: 11 real failures, and they are #516's class
Run in four chunks (~3,171 passed). Every failure was checked against a CPython control run of the identical chunk, because chunking changes test order:
- 2 failures were chunking artifacts —
test_worker_declarationfails the same way on CPython with this ordering. Not PyPy. - 1 is
test_len_returns_int, which CLAUDE.md already documents as a subprocess-timeout flake on this box. - 1 is genuine and cosmetic:
test_expecting_property_name_reason. PyPy'sjsonraises different text, so Nodus's error translation falls through toinvalid JSON at line 1 column 2instead ofexpected property name. Our user-facing JSON errors are coupled to CPython's internals. - 11 are genuine and structural. CPython's control for that chunk was fully green (782 passed). They cluster in
test_nodus_workflow_framework.pyandtest_resume_topology_validation.py, and the cause is:
PermissionError: [WinError 32] The process cannot access the file because it is being used by another process: '...\tmp...\demo.nd'A handle held past its scope. CPython's refcounting closes it the instant the last reference drops; PyPy defers to GC and the tempdir cleanup then fails. This is exactly #516's class — "the store depends on CPython refcounting" — in a different place. #516 fixed one instance; the class has more.
The recommendation
Do not adopt PyPy, and do not keep it as a listed fix direction in its current form. It is slower for the CLI, slower for short programs, ~2x for long ones, and the suite is red.
What is worth keeping is the side effect: PyPy is a working detector for resource leaks that CPython's refcounting hides. Those 11 failures are latent defects on CPython too — files held open past their scope — which matter under load and on any non-refcounting implementation. That is a real finding and it is filed separately rather than buried here.
This issue stays open. The limitation is unchanged and directions 2 (native dispatch loop) and 3 (selective compilation) are untouched; direction 1 is now measured and should be struck.
- 2 failures were chunking artifacts —
Correction to the comment above: PyPy's ceiling is ~15x, not ~2x — we regressed it
My previous comment concluded that PyPy "does not deliver" and that the ~23x note "does not reproduce". The second half was wrong, and it made the first half misleading. Filed as #702.
The ~23x figure reproduces on older trees. Same probe, same machine, same session:
tree CPython PyPy v5.1.0 0.31–0.74M 6.6–8.7M v5.4.0 — 15.75M v5.5.0 — 15.58M v5.6.0 onward 0.65–0.83M 1.4–1.9M d18fc60(#488) is the commit. CPython across the same boundary: 0.73M → 0.65M, inside this box's noise. It is a PyPy-only ~9x regression that no CPython benchmark can see, and it landed in a release with green gates because there is no PyPy in CI.So the corrected reading of this issue's fix direction 1 is not "measured and should be struck" — it is "was working, beating this issue's own 5–10x prediction, and broke in v5.6.0." I struck it on the strength of numbers taken only from
main. Measuring one tree and generalising to a direction was the error; the earlier note was right and I should have bisected before contradicting it.What stands from the previous comment
Independent of the regression, and unchanged:
- The CLI case is still bad.
nodus run hello.ndcosts 1,891 ms on CPython and 4,125 ms on PyPy. PyPy's own interpreter starts faster (104 ms vs 184 ms); the penalty is importing 77 Nodus modules, code that runs once and never warms. A 15x hot loop does not help a script that finishes before the JIT engages, andnodus run's default deadline is 200 ms. - The suite is not green — 11 real failures from handles held past their scope, verified against a fully green CPython control of the identical chunk. Filed as Files are held open past their scope: 11 tests fail on any non-refcounting interpreter (#516's class, new instances) #701.
tzdatais a declared dependency absent fromrequirements.txt; Nodus will not import without it.
Net
PyPy is still not adoptable today — the CLI penalty and #701 are both real. But the reason is no longer "PyPy cannot deliver for this workload." It can, or could: ~15x on the hot loop, which is what this issue wanted from direction 1.
The order of work implied by that: #702 first (recover what we had, and learn why the dispatch loop sits on a cliff), then #701, then reassess adoption. Direction 2 (native dispatch loop) should stay parked until #702's diagnosis lands — if a single instance attribute can cost 9x, that is evidence about the loop's shape which is worth having before investing in rewriting it.
- The CLI case is still bad.
Triage at 5.9.0 — the numbers are stale, and they disagree with each other
This issue's own title and body do not match. The title says "~320-400K instr/sec";
the body says "The throughput ceiling is ~200K Nodus instructions per second" and
"Estimated from benchmark.nd: ~200,000 instructions/second". Both were written on
2026-06-07, against v4.0.0. Ten minor releases have shipped since, so neither figure
should be quoted by anyone until it is re-measured.Fix direction 1 has had real work. "PyPy backend — drop-in for CPython, can 5–10x
throughput on dispatch-heavy loops; no code changes needed" turned out to need code
changes: PyPy has a hard 80-instance-attribute cliff, the VM sat at 79, and #488 crossed
it — costing roughly 9x for three releases before #702 fixed it and recovered ~9.6x.Two durable outcomes from that, both worth having here:
tests/test_vm_attribute_budget.pyfails if a bareVMor aNodusRuntime-built VM
reaches the cliff. It is the only thing in the tree that can see it. Do not raise the
number — shed an attribute.- PyPy is still not adoptable, and not for dispatch reasons: CLI import cost dominates.
So "drop-in, no code changes" is refuted twice over.
Before any further work here, the first commit is a re-measurement on 5.9.0 that replaces
both figures with one, and says which build and which program produced it.- changed the title
[-]LIMITS-001: runtime throughput is the bootstrapping blocker — ~320-400K instr/sec, no JIT[/-][+]LIMITS-001: runtime throughput is the bootstrapping blocker — ~650-850K instr/sec, no JIT[/+]on Sep 6, 2026 Re-baselined 2026-09-06 — the numbers in this issue were stale, and one blocker is gone
Throughput is 650–850K instr/sec, not ~200K. Measured with the VM's own
instruction counter rather than an estimate per loop iteration, CPython 3.11.9:n= 50,000 1.281s 39.0K iter/s 663.8K instr/s (17.0 instr/iter) n= 200,000 5.217s 38.3K iter/s 651.7K instr/s (17.0 instr/iter) n=1,000,000 20.020s 49.9K iter/s 849.2K instr/s (17.0 instr/iter)The title said 320–400K and the body ~200K; both are superseded.
17.0 instr/iteris counted, not assumed — my own first pass at this estimated "about
6 per iteration" and was wrong by nearly 3x, which is why the counter is used.Do not transcribe those numbers either.
tools/benchmark_runtime.pyreports
throughput and startup, so the figure can be re-derived instead of quoted. That
is the whole reason it exists.The PyPy path's largest blocker is fixed
The recorded blocker was CLI import cost — PyPy's interpreter starts faster
than CPython's (104 ms vs 184 ms), and what made it unadoptable was importing
Nodus's own module tree, run-once code the JIT never warms.nodus.cli.cliwas importingnodus.services.serverat module scope for four
commands, which pulls in FastAPI, uvicorn and pydantic — so everynodus run,
nodus fmt,nodus checkandnodus --versionpaid for a web server nobody
asked for. Made lazy in #777:before after import nodus.cli.cli1435 ms 652 ms 2.20x nodus --version1592 ms 640 ms 2.49x nodus run hello.nd1711 ms 1011 ms 1.69x ~700 ms off every CLI invocation. This does not make PyPy adoptable on its
own — it removes the largest single term, and the remaining startup cost is
Nodus's own modules, which is the same shape one level down.What is unchanged
The dispatch ceiling itself. Startup is not throughput, and none of the three
fix directions (PyPy backend, native VM loop, selective compilation) has been
attempted. The measured PyPy figures from the #702 work still stand: 1.6–1.9M
instr/sec onmainafter the 80-attribute cliff was fixed, against CPython's
0.65–0.85M.
Summary
The Nodus VM is implemented as a Python dispatch loop (vm.py). The throughput ceiling is ~200K Nodus instructions per second on a typical developer machine — bounded by CPython's own dispatch overhead, not by the Nodus logic itself. There is no JIT, no loop-invariant hoisting, and no escape from this ceiling within v4.x.
Impact
Observed bound
Estimated from benchmark.nd: ~200,000 instructions/second hot-loop throughput. At that rate, the 10M step limit (embedded default) fires after ~50 seconds of pure compute.
Fix direction
Three paths, all deferred:
None of these are planned for v4.x.
Workaround
For compute-heavy workloads:
nodus run --time-limit Nfor CLI (extends deadline)NodusRuntime(timeout_ms=None, max_steps=None)for embedded (removes limits)subprocess_run()or a host Python functionAffected versions
v4.0.0 (current when filed). Noted in System Limits Audit Check 1.