Skip to content

engine: lazy IF evaluation — skip the untaken branch in the VM #711

Description

@bpowers

Context

Identified during the simlin-engine VM performance round 2 (see docs/design/engine-performance.md "Round 2 wins (2026-06-03)" and the unprioritized run-side swings at the bottom of that file). Sibling perf issues from the same campaign: #601 (threaded dispatch), #602 (lookup hot path), #603 (stride-step flat_offset), #604 (bundle marginal branch/dispatch reductions).

Problem

The bytecode lowers IF THEN ELSE(c, t, f) to "evaluate t; evaluate f; evaluate c; SetCond; If" — both branches are evaluated every timestep and one result is discarded. On C-LEARN there are 3,046 If sites which, counting their SetCond condition chains, are roughly 10% of opcodes in the per-step flow program. Every step pays for the work of the branch that is not selected.

Skipping the untaken branch is value-preserving: the discarded branch cannot affect the selected result (no observable side effects in the VM's expression evaluation), so this is a pure work reduction, not a semantics change.

Why it matters

Performance: ~10% of the per-step opcode budget on the largest model is spent computing values that are immediately thrown away. The run is throughput-bound on the round-2 machine (IPC ~4.5), so cutting executed work per step is the lever.

Scope / why it's a substantial design effort, not a quick patch

Lazy evaluation requires forward-jump opcodes, which touches several subsystems that today all assume straight-line code:

  • codegen — the Expr::If arm in src/simlin-engine/src/compiler/codegen.rs must emit a conditional forward jump over the untaken branch instead of eagerly lowering both arms plus SetCond/If.
  • max_stack_depth validation — the straight-line stack-depth computation in bytecode.rs assumes a single linear path; divergent stack depths across branches (the two arms push, then one is skipped) break that invariant and the validator needs to reason about join points.
  • peephole / fusion jump-target guards — peephole_optimize and fuse_three_address must not fuse across, or relocate, a jump target; they need jump-aware boundaries.
  • symbolic layer — must understand jumps so CompiledSimulation stays symbolizable/disassemblable.
  • wasmgen parity — wasmgen lowers SetCond/If eagerly too, so it needs the matching control-flow lowering (br_if/blocks) to stay behaviorally identical to the VM.

Possible approach

Add forward-jump opcodes (conditional + unconditional), lower Expr::If to "eval c; jump-if-false over t; eval t; jump over f; eval f". Extend max_stack_depth to validate per-branch and at join points. Teach peephole/fusion and the symbolic layer about jump boundaries. Mirror the control flow in wasmgen. Gate the whole thing behind the existing simulate-suite cross-simulator tolerance plus VM-vs-wasm parity tests.

Related context / cautionary note

A provably-equivalent rewrite of vector_elm_map's strict-slice base as a precomputed affine dot product (structurally less work per element) measured a consistent ~5 ms REGRESSION on the 137 ms C-LEARN run — codegen perturbation of the giant inlined eval_bytecode swamped the algorithmic win. Documented in docs/design/engine-performance.md as a negative result reinforcing #604's "measure everything near the eval loop" rule. Any change of this size to the eval loop's control flow must be measured, not argued: the structural win may not survive contact with the inliner.

Refs

  • docs/design/engine-performance.md — "Round 2 wins (2026-06-03)" and the "Lazy If" bullet under the unprioritized run-side swings.
  • src/simlin-engine/src/compiler/codegen.rs — Expr::If lowering.
  • src/simlin-engine/src/bytecode.rs — max_stack_depth validation, peephole_optimize, fuse_three_address.
  • src/simlin-engine/src/vm.rs — SetCond/If handlers.
  • wasmgen — eager SetCond/If lowering needing parity.

Activity

  1. bpowers commented on Jun 4, 2026

    @bpowers
    OwnerAuthor

    Measured NO-GO for now -- census data in docs/design/engine-performance.md (round 3, 2026-06-04), summarized:

    A dispatch census (temporary eval_bytecode counters + a stack-effect branch-span reconstruction over the fused stream the VM actually executes) on the engine-vm-perf branch:

    C-LEARN WORLD3-03
    executed dispatches/step (exact) 30,524 1,112
    flow If sites / executed Ifs per step 1,679 / 1,873 24 / 24
    dispatches lazy-IF would skip 4,859 (15.9%) 36 (3.25%)
    instruction-share estimate of the skip ~1.5% (~35k of ~2.4M instr/step) trivial

    The 15.9% dispatch share collapses to ~1.5% instructions because 93% of the skippable opcodes are cheap scalar loads/binops (LoadVar 30%, BinStackVar 11%, LoadConstant 11%, ...; only 3.7% Lookup + 2.9% Apply), while per-step instruction count is dominated by per-element array/lookup/module work lazy-IF never touches. ~1.5% is below the ~4% build-layout measurement floor documented in the perf doc (negative result #3's methodology note), and this is the most cross-cutting change of the round-3 candidates (forward jumps touch the stack-depth validator, peephole/fusion jump maps, the symbolic layer, and wasmgen). The lowering also adds two jump dispatches per executed If back.

    Two findings worth keeping:

    • 69.8% of the skippable dispatches sit behind constant conditions (1,300 of 1,679 flow If sites take the same branch for the entire run -- policy switches). Those are compile-time / engine: time-invariant variable hoisting — a constant phase evaluated once per run_to #712-family targets (branch specialization), not runtime-jump targets.
    • Exact dispatch counts: C-LEARN executes ~30.5k dispatches/step, far below earlier loose estimates -- useful calibration for sizing any future dispatch-level work.

    Revisit conditions: a mispredict-bound microarchitecture (the round-1 Ryzen profile, where the dispatch indirect branch dominates and the If/SetCond data-dependent selection has real misprediction cost), or as part of a threaded-dispatch rewrite (#601) where the control-flow structure changes anyway.

  2. added
    engineIssues with the rust-based simulation engine
    enhancementNew feature or request
    on Jun 8, 2026
  3. bpowers commented on Aug 10, 2026

    @bpowers
    OwnerAuthor

    Correction to the measurement premise in this issue's NO-GO — the verdict is unchanged

    The decision stands. Only its stated reason is wrong, and the wrong reason is the kind that shuts down further thought, so it is worth correcting where people will read it.

    docs/design/engine-performance.md records the NO-GO as:

    the instruction share is only ~1.5% (~35k of ~2.4M instr/step): below the ~4% layout-noise measurement floor

    That applies a cycles/binary-layout noise floor to an instruction-count measurement. The two floors are three orders of magnitude apart.

    Measured on the C-LEARN run across six independent build+run pairs of identical source (Ryzen 9950X):

    channel sd across builds
    retired instructions 0.026%
    branches 0.028%
    cycles, quiet machine 1.65%
    cycles, machine under load 9.9%–11%

    A null control makes the same point directly — the identical binary run as both sides of an interleaved A/B, 5 rounds, medians:

    instructions   -0.003%
    branches       -0.004%
    cycles         -1.540%     <- an apparent "win" from nothing
    

    Retired instructions are a property of the program; cycles are a property of the machine executing it. So ~1.5% of instructions is roughly 58 sigma on that channel — comfortably measurable, not below any floor. The ~4% figure is real, but it bounds wall-clock/cycles claims from a single build pair, not instruction counts.

    What the verdict actually rests on

    Two reasons survive, and they are sufficient:

    1. Small absolute win for the highest design cost of the three candidates. Lazy If needs forward-jump opcodes, which touch codegen, max_stack_depth join validation, the peephole/fusion jump maps, the symbolic layer, and wasmgen parity. ~1.5% of instructions does not pay for that.
    2. The cheap part of the win is already taken. Fusing SetCond;If[;AssignCurr] into conditional-select opcodes removes 12.0% of executed dispatches against this issue's projected 15.9%, with none of that machinery — no jumps, no stack-depth join reasoning, no symbolic or wasmgen changes. It rests on the pair being adjacent by construction (compiler::codegen's Expr::If arm is the sole producer of both and emits them together; measured executed counts are exactly equal at 1,874,169 each on C-LEARN).

    So the remaining upside here is the residual after that fusion, against the full forward-jump design cost.

    Why this is worth a comment rather than a silent doc edit

    A wrong measurement premise in a design doc does not get caught by review: reviewers check the reasoning against the stated premise, not the premise against the world. "Below the measurement floor" reads as a fact and ends the discussion. This one had already carried a verdict, which is the evidence it was doing damage rather than sitting inert.

    Anyone picking this issue up would have read the invalid floor claim here first, so fixing only the doc would have preserved the trap in the place most likely to be read.

    The general convention that would have prevented it: when a doc or issue records a measurement-derived verdict, state which channel the number came from. A bare "only ~1.5%" invites comparison against whatever floor the reader has in mind.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    engineIssues with the rust-based simulation engineenhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions