Skip to content

LIMITS-001: runtime throughput is the bootstrapping blocker — ~650-850K instr/sec, no JIT #173

Description

@Masterplanner25

Summary

The Nodus VM is implemented as a Python dispatch loop (vm.py). The throughput ceiling is ~200K Nodus instructions per second on a typical developer machine — bounded by CPython's own dispatch overhead, not by the Nodus logic itself. There is no JIT, no loop-invariant hoisting, and no escape from this ceiling within v4.x.

Impact

  • CLI default: 200ms wall-clock → ~40,000 compute instructions before timeout fires
  • Tight computation-heavy programs (sorting large lists, numerical algorithms) hit the timeout before producing meaningful output
  • The ceiling cannot be raised by configuration — only by running fewer instructions or using a faster runtime

Observed bound

Estimated from benchmark.nd: ~200,000 instructions/second hot-loop throughput. At that rate, the 10M step limit (embedded default) fires after ~50 seconds of pure compute.

Fix direction

Three paths, all deferred:

  1. PyPy backend — drop-in for CPython, can 5–10x throughput on dispatch-heavy loops; no code changes needed
  2. Native VM loop — rewrite the dispatch loop in Rust/C as a CPython extension; PyO3 scaffolding already exists in nodus-native-memory-engine
  3. Selective compilation — compile hot paths to Python bytecode or native code; requires significant IR work

None of these are planned for v4.x.

Workaround

For compute-heavy workloads:

  • nodus run --time-limit N for CLI (extends deadline)
  • NodusRuntime(timeout_ms=None, max_steps=None) for embedded (removes limits)
  • Delegate heavy computation to subprocess_run() or a host Python function

Affected versions

v4.0.0 (current when filed). Noted in System Limits Audit Check 1.

Activity

  1. added this to the v5.0 milestone on Jun 8, 2026
  2. Masterplanner25 commented on Aug 17, 2026

    @Masterplanner25
    OwnerAuthor

    Triage 2026-08-17 (against v5.0.2). The measurement in this issue is wrong — the ceiling is roughly 2.3x higher
    than stated.
    The conclusion still holds; the number should not be quoted.

    Measured on 5.0.2, a tight arithmetic loop through the real embedding API:

    5,100,010 instructions in 11.24 s  ->  ~454,000 instr/sec
    

    against the ~200K instr/sec this issue claims. Measured on a machine that is
    currently slower than normal (see the timing note in CLAUDE.md), so the true
    figure is likely higher still.

    Everything structural in the issue remains true: the VM is a CPython dispatch loop,
    there is no JIT, no loop-invariant hoisting, no escape analysis. Only the throughput
    figure needs correcting, and it is the kind of number that gets quoted in
    positioning documents, which is why it is worth fixing rather than leaving.

  3. Masterplanner25 commented on Aug 19, 2026

    @Masterplanner25
    OwnerAuthor

    Refiling this as the bootstrapping blocker

    Re-measured on 5.0.4, and reframing what this issue is. It has been carried as a
    generic performance note at severity:low; it is in fact the thing standing
    between Nodus and self-hosting, which docs/language/LANGUAGE_VISION.md now
    states as a decided direction rather than an aspiration.

    Current measurement

    Executed rather than estimated, embedded runtime with no step or time limit:

    SRC = '''
    fn main() {
        let i = 0i
        let acc = 0i
        while (i < 1000000i) { acc = acc + i; i = i + 1i }
        print("acc=\(acc)")
    }
    '''
    rt = NodusRuntime(timeout_ms=None, max_steps=None)
    acc=499999500000
    1,000,000 loop iterations in 54.13s
    stats: {'instructions_executed': 17000021, 'coroutines_spawned': 0}
    

    ≈ 314,000 instructions/sec — about 17 instructions per trivial loop
    iteration. Somewhat better than this issue's original ~200K estimate, and the
    same order.

    Why this is the bootstrapping blocker specifically

    The three other prerequisites the roadmap names are in good shape:

    Prerequisite State
    Stable language semantics core surfaces are Stable; 18 adversarial corpus reads produced zero findings against lexer/parser/VM/types
    Stable bytecode instruction set 49 opcodes, BYTECODE_VERSION 4, frozen since v1.0 and gate-enforced
    Reliable module system import/export stable; #158 (no privacy) is an annoyance, not a barrier
    Sufficiently expressive stdlib the weak one — no string slice, no byte type (#170), no closure upvalue mutation (#156)

    And the shape of the task is already demonstrated: examples/expr_compiler.nd
    is a character-level lexer, recursive-descent parser and tree-walking evaluator
    written entirely in Nodus, 231 lines.

    So the gap is not can the language express a compiler — it demonstrably can.
    The gap is arithmetic. The pipeline that would need rewriting is ~2,300 lines of
    lexer/parser/AST plus ~1,900 lines of compiler. A Nodus lexer without a slice
    operator indexes character by character; at ~314K instr/sec, compiling a single
    mid-sized source file lands in the seconds, and a self-hosted compiler compiling
    itself is the workload that has to be affordable.

    Nothing in the other three prerequisites changes that. Only this does.

    On the fix directions

    The three listed still look right, and the middle one is better positioned than
    when this was filed:

    1. PyPy backend — no code changes, 5–10× on dispatch-heavy loops. Cheapest
      thing to actually measure, and worth measuring before choosing anything else:
      if it lands near 3M instr/sec the calculus changes completely.
    2. Native VM loop — PyO3 scaffolding exists and is now proven in production,
      not just present: nodus-native-memory-engine v0.1.1 ships a Rust extension
      with a pure-Python fallback. The pattern of "native where it pays, Python
      fallback everywhere" is established in this ecosystem rather than theoretical.
    3. Selective compilation — still the largest IR investment, still last.

    Suggested first step is (1) as a measurement rather than a commitment: run the
    same 1M-iteration probe under PyPy and record the number. It is an afternoon,
    and it decides whether (2) is necessary or merely nice.

    What changes on this issue

    • severity:low → severity:medium. Not high: nothing is broken, no user is
      blocked today, and the workarounds listed are real. But "low" reads as this
      does not matter
      , and it is the single blocker on a stated direction of the
      language.
    • Adding type:design-question, because choosing among the three paths is a
      decision, not an implementation task.
    • The "None of these are planned for v4.x" line is now just true-and-stale; v4.x
      is three majors back.

    Provenance

    Surfaced while auditing how visible the bootstrapping goal is and how close the
    language actually is to it — the answer being that it was documented nowhere a
    reader would find it, and that this issue is what stands in the way. Docs fixed
    separately; this is the tracking half.

  4. added
    severity:mediumUX broken but workaround exists
    type:design-questionLanguage-level design decision, not a simple fix
    and removed on Aug 19, 2026
  5. changed the title [-]LIMITS-001: Python interpreter throughput ceiling — ~200K instr/sec, no JIT[/-] [+]LIMITS-001: runtime throughput is the bootstrapping blocker — ~314K instr/sec, no JIT[/+] on Aug 19, 2026
  6. Masterplanner25 commented on Aug 20, 2026

    @Masterplanner25
    OwnerAuthor

    Measured: PyPy is ~23× faster, and Nodus runs on it unmodified

    This issue's first fix direction was "PyPy backend — no code changes, 5–10× on
    dispatch-heavy loops"
    , and my earlier comment argued it should be measured
    before
    committing to a native VM, because if it landed near 3M instr/sec the
    calculus changes completely. It lands well past that.

    The numbers

    Same 1M-iteration probe (17,000,021 instructions), same machine, three trials
    each, best and median reported because a single sample on this box is not
    trustworthy — see the correction below.

    runtime best median vs CPython 3.11
    CPython 3.11.9 407,097 401,694 —
    CPython 3.14.2 454,600 383,872 ~1× (noise)
    PyPy 7.3.23 (3.11.15) 10,048,618 9,384,307 ~23×

    PyPy trial 1 was 3.06s against 1.69s and 1.81s for trials 2 and 3 — JIT warmup,
    so the warm figure is the honest one for a long-running compile.

    Verified as real work rather than skipped work: PyPy produces
    acc=499999500000 and executes the same 17,000,021 instructions.

    Correcting my own figure

    My earlier comment reported ~314K instr/sec from a single run, and that number
    went into LANGUAGE_VISION.md and the README. Re-measuring gives 365K–454K on
    CPython across six trials. The 314K was a slow sample on a noisy box, not a
    wrong method — but it was one sample presented as a measurement, which is the
    thing I should not have done. The docs should say ~400K, and I will correct
    them.

    It does not change any conclusion. The ratio is what matters and it is ~23× either
    way.

    CPython 3.14 buys nothing

    Worth recording so nobody repeats it: 3.14 is within noise of 3.11 on this
    workload. Upgrading the interpreter is not a free win here.

    Nodus needs no changes to run on PyPy

    The only dependency is tzdata, which is pure data. pip install tzdata, set
    PYTHONPATH, and run_source works first try.

    One thing blocks it, and it is a real bug either way

    A 65-test subset (task graph, workflow DSL, state policies, coroutines, channels)
    gives 9 failed, 65 passed — and all nine share a single root cause:

    cannot commit transaction - SQL statements in progress
    

    SQLiteWorkflowStore never closes the cursors conn.execute(...) returns. On
    CPython refcounting frees them immediately; PyPy's GC does not, so the statement
    is still open at commit. Filed as #516.

    That is not a PyPy incompatibility to work around — it is the store depending on
    when CPython happens to free an object for correctness, which is a latent
    defect on CPython too. Coroutines, channels, task graph and state policies all
    pass unmodified.

    What this means for the three fix directions

    1. PyPy backend — now the obvious first move. ~23×, no code changes, one
      blocking bug already identified and small. It does not need to become the
      supported runtime to be worth having: adding it to CI would catch the whole
      class of refcount-dependent resource bugs that CPython structurally cannot.
    2. Native VM loop — should wait. It is the largest investment on the list and
      PyPy may make it unnecessary. Reassess after SQLiteWorkflowStore leaks cursors: the store depends on CPython refcounting to commit #516 and a full suite run.
    3. Selective compilation — unchanged, still last.

    What this means for bootstrapping

    At ~400K instr/sec a self-hosted compiler compiling itself was the blocker this
    issue describes. At ~9.4M it is a different conversation — not solved, but no
    longer the thing standing in the way. The next honest step is a real workload
    rather than a synthetic loop: run examples/expr_compiler.nd over a large input
    under both runtimes and compare. A tight arithmetic loop is the case a JIT is
    best at, so 23× is an upper bound, not a promise.

    Method

    # 33 MB, extracted, no install
    pypy.exe -m pip install tzdata
    $env:PYTHONPATH="C:/dev/Coding Language/src"
    pypy.exe bench.py      # 3 trials, NodusRuntime(timeout_ms=None, max_steps=None)

    The probe is the one from my earlier comment, with the print removed so output
    handling is not in the measurement and with three trials instead of one.

  7. 2 remaining items

  8. Masterplanner25 commented on Aug 20, 2026

    @Masterplanner25
    OwnerAuthor

    The compiler workload: PyPy is ~4–5×, not ~23×

    My last comment measured PyPy at ~23× on the 1M-iteration arithmetic probe and
    concluded that the native VM "should wait" because "PyPy may make it unnecessary".
    That was measured on the workload a JIT is best at. Re-measured on the workload that
    actually matters for bootstrapping, the answer is different enough to change the
    conclusion.

    The workload

    examples/expr_compiler.nd — the character-level lexer, recursive-descent parser and
    tree-walking evaluator this issue already cites as proof the language can express a
    compiler — driven over a non-trivial expression n times:

    ((1 + 2) * (3 + 4)) - -5 / 2 + 3.5 * 2 * (10 - 4 / 2)
    

    8,209 VM instructions per iteration. In-process via run_source, so interpreter
    startup is out of the number; three trials per cell, best reported.

    The numbers

    n instructions CPython 3.11.9 PyPy 7.3.23 (3.11.15) ratio
    25 205,215 323,764/s 324,095/s 1.00×
    100 820,740 337,151/s 1,388,545/s 4.12×
    400 3,282,840 319,053/s 1,741,086/s 5.46×
    1600 13,131,240 209,252/s 856,436/s 4.09×

    Same thing through the CLI as whole processes, which folds in startup and is the
    number a user would feel:

    n CPython wall PyPy wall ratio (startup subtracted)
    50 5.10 s 3.39 s 1.81×
    200 10.17 s 4.25 s 3.14×
    800 36.58 s 8.98 s 4.70×

    Startup itself: CPython 1.88 s, PyPy 1.60 s.

    PyPy peaks around 1.7M instr/sec here, against 10.05M on the arithmetic probe —
    5.8× lower on the same machine, same interpreter, same day. The JIT needs a hot loop;
    a recursive-descent parser is many cold-ish paths, and it warms slowly: at n=25 PyPy
    has no advantage at all, and it is still climbing at n=400.

    Do not read the n=1600 row as a decline. That row is box noise — a later run of
    the same n=1600 case on CPython gave 324,640 instr/sec, in line with the plateau,
    and the trial spread on that row was 62.75 / 127.15 / 69.01 s. This machine's timing
    instability is documented in CLAUDE.md and it is fully live. The honest reading of
    the table is: CPython ~320K instr/sec flat, PyPy ~1.7M once warm, ratio 4–5.5×.

    A third of CPython's throughput is not dispatch overhead

    While chasing that apparent instability I found something that changes the framing of
    this issue. Filed as #522.

    This issue says the ceiling is "bounded by CPython's own dispatch overhead, not by the
    Nodus logic itself". A third of it is neither: the VM emits and retains an event
    object for every function call, every return, and every 100 instructions, into an
    unbounded list that nothing reads unless a host attaches a sink.

    Same source, same machine, only RuntimeEventBus.emit stubbed to a no-op:

    n bus on bus off speedup events retained
    400 302,679/s 454,941/s 1.50× 344,035
    1600 324,640/s 490,539/s 1.51× 1,376,119

    tracemalloc at n=400: 131.4 MB live, 74.4 MB of it the event log — 57% of
    everything the run allocates, ~23 bytes per instruction executed.

    So there is a 1.5× available with no new runtime, no JIT and no native code, and
    the memory side is arguably the bigger deal for bootstrapping: at 23 bytes per
    instruction, a self-hosted compiler executing a billion instructions retains tens of
    gigabytes of vm_call records. Throughput you can wait out; that you cannot.

    What this does to the three fix directions

    1. PyPy backend — still worth having, but on the honest workload it is ~4–5×,
      not 23×
      , landing near 1.7M instr/sec. My "if it lands near 3M the calculus
      changes completely" test is not met by the compiler workload; it was met only by
      the arithmetic probe. It remains the cheapest real gain, it needs no code changes,
      and adding it to CI catches the refcount-dependent resource bugs CPython cannot
      (which is how SQLiteWorkflowStore leaks cursors: the store depends on CPython refcounting to commit #516 was found). But it does not retire the question.
    2. Native VM loop — I retract "should wait", at least as a conclusion drawn
      from the PyPy number. 4–5× on a compiler workload does not make a native dispatch
      loop unnecessary. It should be reassessed on its own merits, after The VM retains an event per function call and return, unbounded: 1.5x throughput and 74 MB per 3M instructions, with no consumer #522, not
      dismissed on the strength of a benchmark that flattered the alternative.
    3. Selective compilation — unchanged, still last.

    And a fourth that was not on the list, which should now be first because it is
    strictly cheaper than all three: #522 — stop retaining telemetry nobody reads.
    1.5× and the memory wall, in ordinary Python, in the existing runtime.

    Correcting myself, again

    Two things I got wrong on this issue and one method note.

    • I reported ~314K instr/sec from a single run; re-measurement gave 365K–454K, and
      the docs were corrected to ~400K in docs: re-measure runtime throughput, and record the PyPy result #517. The compiler workload measured here sits
      at ~320K, i.e. lower than the arithmetic figure — which is expected and worth
      keeping distinct in the docs: the headline number is a hot arithmetic loop, and
      real programs are slower.
    • I concluded the native VM should wait, on a 23× that does not survive contact with
      the workload this issue is about. Retracted above.

    The method note: the first attempt at this benchmark ran for over ten minutes on a
    workload the CLI finished in one second, because the harness passed
    filename="compiler_bench.nd" to run_source — and run_source executes the file
    named by filename and discards the source string it was given
    . It was running the
    on-disk copy at 3200 iterations instead of the 25-iteration string it was handed.
    That is #521, present since v0.4.0, and I would not have found it without this
    benchmark being wrong first.

    Provenance

    Requested follow-up to the PyPy measurement above: run the same comparison on
    expr_compiler.nd rather than a synthetic loop, on the reasoning that a tight
    arithmetic loop is a JIT's best case and therefore an upper bound rather than a
    promise. It was.

  9. changed the title [-]LIMITS-001: runtime throughput is the bootstrapping blocker — ~314K instr/sec, no JIT[/-] [+]LIMITS-001: runtime throughput is the bootstrapping blocker — ~320-400K instr/sec, no JIT[/+] on Aug 20, 2026
  10. Masterplanner25 commented on Aug 23, 2026

    @Masterplanner25
    OwnerAuthor

    #522 is fixed, and it moves this issue's baseline. Measured on
    examples/expr_compiler.nd (400 expressions, 1.96M instructions), same machine,
    before and after the change:

    before after
    throughput 227,401 instr/sec 477,417 instr/sec
    events retained 206,382 0
    live memory 80.2 MB 0.4 MB

    That is 2.10×, on CPython, with no JIT and no native VM — because a third to
    a half of the time was the runtime building vm_call / vm_return event objects
    that nothing read, and the guard now sits before the object is constructed rather
    than after.

    Two consequences for this issue as written:

    1. The "~320-400K instr/sec" figure in the title is now the floor, not the
      ceiling.
      A larger run (7.8M instructions) measures 354K/s post-fix on a
      loaded box; the 400-expression run measures 477K/s. Re-measure before quoting
      a ceiling.

    2. The attribution needs revising. This issue says the limit is "CPython's own
      dispatch overhead, not the Nodus logic itself". Half of it was neither — it was
      telemetry with no consumer. Whatever remains should be re-profiled before it is
      used to argue for PyPy or a native VM, for the same reason the PyPy figure was
      retracted earlier in this issue: a benchmark that flatters the alternative.

    The memory side matters more for the bootstrapping argument here. At ~23 bytes per
    instruction retained, a self-hosted compiler executing a billion instructions was
    accumulating tens of gigabytes of vm_call records. That wall is gone rather than
    raised — retention no longer scales with the run at all.

  11. removed this from the v5.0 milestone on Aug 25, 2026
  12. Masterplanner25 commented on Aug 31, 2026

    @Masterplanner25
    OwnerAuthor

    Re-measured at 5.8.0: the number in this issue is stale by 3–5x

    Filed against v4.0.0. Eight minor releases later the ceiling has moved, and nothing recorded it — so the figure people plan around is wrong in the direction that makes the problem look worse than it is.

    The title and body already disagreed (~320-400K vs ~200,000), which is its own signal that neither had been re-derived.

    Measured

    CPython 3.11.9, current main, NodusRuntime(timeout_ms=None, max_steps=None), best-of-N, instructions_executed from get_execution_stats() divided by wall clock:

    shape instr/sec
    tight arithmetic loop (400K iterations) 1,125K
    tight arithmetic loop (200K iterations) 935K
    function calls (add(a, b) in a loop) 970K
    list list_push then indexed read 950K

    ~0.9–1.1M instructions/second, consistent across three unrelated program shapes rather than one favourable microbenchmark.

    What that changes in this issue's own impact section

    • "CLI default: 200ms wall-clock → ~40,000 compute instructions before timeout fires" — EXECUTION_TIMEOUT_MS is still 200, but 200 ms now buys roughly 190,000 instructions, not 40,000.
    • "the 10M step limit (embedded default) fires after ~50 seconds of pure compute" — closer to 10 seconds.

    What has not changed

    The issue's actual claim stands: it is still a CPython dispatch loop, there is still no JIT, and throughput is still the bootstrapping blocker. A ~5x improvement does not change which of the three fix directions apply.

    One of them has since been measured and is worth recording here rather than leaving in a session note: PyPy runs Nodus roughly 23x faster unmodified — the "drop-in, no code changes" path this issue lists first. It is not adoptable today because #516 (a cursor leak) blocks the suite under PyPy. CPython 3.14 was also tried and buys nothing.

    Staying open: the limitation is real, only the number was wrong. Re-derive it rather than trusting this comment in another eight releases — the script is four lines against get_execution_stats().

  13. Masterplanner25 commented on Aug 31, 2026

    @Masterplanner25
    OwnerAuthor

    PyPy measured end to end — it does not deliver, and the ~23x figure does not reproduce

    #516 (the cursor leak that blocked this) closed 2026-08-20, so I stood PyPy up and measured rather than inheriting the number. PyPy 7.3.23 / Python 3.11.15 (matching the project's CPython 3.11.9), official win64 build, .venv-pypy, requirements.txt minus mypy (its ast-serialize dependency has no PyPy wheel; irrelevant here since CI type-checks on CPython). tzdata had to be installed separately — it is a declared dependency in pyproject.toml but not in requirements.txt, and time_module.py imports ZoneInfo("UTC") at module import, so Nodus does not start without it.

    Throughput: 1.5–2.9x, and negative for short programs

    Same three shapes, both interpreters, back to back in one window (this box's timing is unstable, so the ratio is worth more than the absolutes):

    shape CPython PyPy
    tight arithmetic loop 638K 1,859K 2.9x
    function calls 883K 1,319K 1.5x
    list index + append 760K 1,438K 1.9x

    Then the same loop at four sizes, which is the more useful cut:

    iterations CPython PyPy
    50,000 782K 375K 0.48x — PyPy is 2x slower
    200,000 598K 821K 1.4x
    1,000,000 653K 1,462K 2.2x
    4,000,000 897K 1,551K 1.7x

    It crosses over between 50K and 200K iterations and plateaus near 1.5M instr/sec. The "~23x" figure carried in session notes does not reproduce; I cannot speak to its provenance, and it should not be quoted again without a fresh measurement. This issue's own fix direction — "PyPy backend — drop-in for CPython, can 5–10x throughput on dispatch-heavy loops; no code changes needed" — is not supported by measurement either.

    The CLI case is worse, and it is not interpreter startup

    CPython PyPy
    bare interpreter (-c pass) 184 ms 104 ms
    nodus run hello.nd 1,891 ms 4,125 ms

    PyPy's own startup is faster. The 2.2x penalty is importing and compiling Nodus's 77 modules — code that runs once and never warms up. For nodus run, whose default deadline is 200 ms, PyPy is strictly worse.

    The suite is not green: 11 real failures, and they are #516's class

    Run in four chunks (~3,171 passed). Every failure was checked against a CPython control run of the identical chunk, because chunking changes test order:

    • 2 failures were chunking artifacts — test_worker_declaration fails the same way on CPython with this ordering. Not PyPy.
    • 1 is test_len_returns_int, which CLAUDE.md already documents as a subprocess-timeout flake on this box.
    • 1 is genuine and cosmetic: test_expecting_property_name_reason. PyPy's json raises different text, so Nodus's error translation falls through to invalid JSON at line 1 column 2 instead of expected property name. Our user-facing JSON errors are coupled to CPython's internals.
    • 11 are genuine and structural. CPython's control for that chunk was fully green (782 passed). They cluster in test_nodus_workflow_framework.py and test_resume_topology_validation.py, and the cause is:
    PermissionError: [WinError 32] The process cannot access the file because
    it is being used by another process: '...\tmp...\demo.nd'
    

    A handle held past its scope. CPython's refcounting closes it the instant the last reference drops; PyPy defers to GC and the tempdir cleanup then fails. This is exactly #516's class — "the store depends on CPython refcounting" — in a different place. #516 fixed one instance; the class has more.

    The recommendation

    Do not adopt PyPy, and do not keep it as a listed fix direction in its current form. It is slower for the CLI, slower for short programs, ~2x for long ones, and the suite is red.

    What is worth keeping is the side effect: PyPy is a working detector for resource leaks that CPython's refcounting hides. Those 11 failures are latent defects on CPython too — files held open past their scope — which matter under load and on any non-refcounting implementation. That is a real finding and it is filed separately rather than buried here.

    This issue stays open. The limitation is unchanged and directions 2 (native dispatch loop) and 3 (selective compilation) are untouched; direction 1 is now measured and should be struck.

  14. Masterplanner25 commented on Aug 31, 2026

    @Masterplanner25
    OwnerAuthor

    Correction to the comment above: PyPy's ceiling is ~15x, not ~2x — we regressed it

    My previous comment concluded that PyPy "does not deliver" and that the ~23x note "does not reproduce". The second half was wrong, and it made the first half misleading. Filed as #702.

    The ~23x figure reproduces on older trees. Same probe, same machine, same session:

    tree CPython PyPy
    v5.1.0 0.31–0.74M 6.6–8.7M
    v5.4.0 — 15.75M
    v5.5.0 — 15.58M
    v5.6.0 onward 0.65–0.83M 1.4–1.9M

    d18fc60 (#488) is the commit. CPython across the same boundary: 0.73M → 0.65M, inside this box's noise. It is a PyPy-only ~9x regression that no CPython benchmark can see, and it landed in a release with green gates because there is no PyPy in CI.

    So the corrected reading of this issue's fix direction 1 is not "measured and should be struck" — it is "was working, beating this issue's own 5–10x prediction, and broke in v5.6.0." I struck it on the strength of numbers taken only from main. Measuring one tree and generalising to a direction was the error; the earlier note was right and I should have bisected before contradicting it.

    What stands from the previous comment

    Independent of the regression, and unchanged:

    • The CLI case is still bad. nodus run hello.nd costs 1,891 ms on CPython and 4,125 ms on PyPy. PyPy's own interpreter starts faster (104 ms vs 184 ms); the penalty is importing 77 Nodus modules, code that runs once and never warms. A 15x hot loop does not help a script that finishes before the JIT engages, and nodus run's default deadline is 200 ms.
    • The suite is not green — 11 real failures from handles held past their scope, verified against a fully green CPython control of the identical chunk. Filed as Files are held open past their scope: 11 tests fail on any non-refcounting interpreter (#516's class, new instances) #701.
    • tzdata is a declared dependency absent from requirements.txt; Nodus will not import without it.

    Net

    PyPy is still not adoptable today — the CLI penalty and #701 are both real. But the reason is no longer "PyPy cannot deliver for this workload." It can, or could: ~15x on the hot loop, which is what this issue wanted from direction 1.

    The order of work implied by that: #702 first (recover what we had, and learn why the dispatch loop sits on a cliff), then #701, then reassess adoption. Direction 2 (native dispatch loop) should stay parked until #702's diagnosis lands — if a single instance attribute can cost 9x, that is evidence about the loop's shape which is worth having before investing in rewriting it.

  15. Masterplanner25 commented on Sep 2, 2026

    @Masterplanner25
    OwnerAuthor

    Triage at 5.9.0 — the numbers are stale, and they disagree with each other

    This issue's own title and body do not match. The title says "~320-400K instr/sec";
    the body says "The throughput ceiling is ~200K Nodus instructions per second" and
    "Estimated from benchmark.nd: ~200,000 instructions/second". Both were written on
    2026-06-07, against v4.0.0. Ten minor releases have shipped since, so neither figure
    should be quoted by anyone until it is re-measured.

    Fix direction 1 has had real work. "PyPy backend — drop-in for CPython, can 5–10x
    throughput on dispatch-heavy loops; no code changes needed"
    turned out to need code
    changes: PyPy has a hard 80-instance-attribute cliff, the VM sat at 79, and #488 crossed
    it — costing roughly 9x for three releases before #702 fixed it and recovered ~9.6x.

    Two durable outcomes from that, both worth having here:

    • tests/test_vm_attribute_budget.py fails if a bare VM or a NodusRuntime-built VM
      reaches the cliff. It is the only thing in the tree that can see it. Do not raise the
      number — shed an attribute.
    • PyPy is still not adoptable, and not for dispatch reasons: CLI import cost dominates.
      So "drop-in, no code changes" is refuted twice over.

    Before any further work here, the first commit is a re-measurement on 5.9.0 that replaces
    both figures with one, and says which build and which program produced it.

  16. changed the title [-]LIMITS-001: runtime throughput is the bootstrapping blocker — ~320-400K instr/sec, no JIT[/-] [+]LIMITS-001: runtime throughput is the bootstrapping blocker — ~650-850K instr/sec, no JIT[/+] on Sep 6, 2026
  17. Masterplanner25 commented on Sep 6, 2026

    @Masterplanner25
    OwnerAuthor

    Re-baselined 2026-09-06 — the numbers in this issue were stale, and one blocker is gone

    Throughput is 650–850K instr/sec, not ~200K. Measured with the VM's own
    instruction counter rather than an estimate per loop iteration, CPython 3.11.9:

      n=   50,000   1.281s   39.0K iter/s   663.8K instr/s   (17.0 instr/iter)
      n=  200,000   5.217s   38.3K iter/s   651.7K instr/s   (17.0 instr/iter)
      n=1,000,000  20.020s   49.9K iter/s   849.2K instr/s   (17.0 instr/iter)
    

    The title said 320–400K and the body ~200K; both are superseded. 17.0 instr/iter is counted, not assumed — my own first pass at this estimated "about
    6 per iteration" and was wrong by nearly 3x, which is why the counter is used.

    Do not transcribe those numbers either. tools/benchmark_runtime.py reports
    throughput and startup, so the figure can be re-derived instead of quoted. That
    is the whole reason it exists.

    The PyPy path's largest blocker is fixed

    The recorded blocker was CLI import cost — PyPy's interpreter starts faster
    than CPython's (104 ms vs 184 ms), and what made it unadoptable was importing
    Nodus's own module tree, run-once code the JIT never warms.

    nodus.cli.cli was importing nodus.services.server at module scope for four
    commands, which pulls in FastAPI, uvicorn and pydantic — so every nodus run,
    nodus fmt, nodus check and nodus --version paid for a web server nobody
    asked for. Made lazy in #777:

    before after
    import nodus.cli.cli 1435 ms 652 ms 2.20x
    nodus --version 1592 ms 640 ms 2.49x
    nodus run hello.nd 1711 ms 1011 ms 1.69x

    ~700 ms off every CLI invocation. This does not make PyPy adoptable on its
    own — it removes the largest single term, and the remaining startup cost is
    Nodus's own modules, which is the same shape one level down.

    What is unchanged

    The dispatch ceiling itself. Startup is not throughput, and none of the three
    fix directions (PyPy backend, native VM loop, selective compilation) has been
    attempted. The measured PyPy figures from the #702 work still stand: 1.6–1.9M
    instr/sec on main after the 80-attribute cliff was fixed, against CPython's
    0.65–0.85M.

  18. added a commit that references this issue on Sep 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions