Repository navigation
engine: lazy IF evaluation — skip the untaken branch in the VM #711
Description
Activity
Measured NO-GO for now -- census data in docs/design/engine-performance.md (round 3, 2026-06-04), summarized:
A dispatch census (temporary eval_bytecode counters + a stack-effect branch-span reconstruction over the fused stream the VM actually executes) on the engine-vm-perf branch:
C-LEARN WORLD3-03 executed dispatches/step (exact) 30,524 1,112 flow If sites / executed Ifs per step 1,679 / 1,873 24 / 24 dispatches lazy-IF would skip 4,859 (15.9%) 36 (3.25%) instruction-share estimate of the skip ~1.5% (~35k of ~2.4M instr/step) trivial The 15.9% dispatch share collapses to ~1.5% instructions because 93% of the skippable opcodes are cheap scalar loads/binops (LoadVar 30%, BinStackVar 11%, LoadConstant 11%, ...; only 3.7% Lookup + 2.9% Apply), while per-step instruction count is dominated by per-element array/lookup/module work lazy-IF never touches. ~1.5% is below the ~4% build-layout measurement floor documented in the perf doc (negative result #3's methodology note), and this is the most cross-cutting change of the round-3 candidates (forward jumps touch the stack-depth validator, peephole/fusion jump maps, the symbolic layer, and wasmgen). The lowering also adds two jump dispatches per executed If back.
Two findings worth keeping:
- 69.8% of the skippable dispatches sit behind constant conditions (1,300 of 1,679 flow If sites take the same branch for the entire run -- policy switches). Those are compile-time / engine: time-invariant variable hoisting — a constant phase evaluated once per run_to #712-family targets (branch specialization), not runtime-jump targets.
- Exact dispatch counts: C-LEARN executes ~30.5k dispatches/step, far below earlier loose estimates -- useful calibration for sizing any future dispatch-level work.
Revisit conditions: a mispredict-bound microarchitecture (the round-1 Ryzen profile, where the dispatch indirect branch dominates and the If/SetCond data-dependent selection has real misprediction cost), or as part of a threaded-dispatch rewrite (#601) where the control-flow structure changes anyway.
- addedengineIssues with the rust-based simulation engineIssues with the rust-based simulation engineenhancementNew feature or requestNew feature or request
on Jun 8, 2026 Correction to the measurement premise in this issue's NO-GO — the verdict is unchanged
The decision stands. Only its stated reason is wrong, and the wrong reason is the kind that shuts down further thought, so it is worth correcting where people will read it.
docs/design/engine-performance.mdrecords the NO-GO as:the instruction share is only ~1.5% (~35k of ~2.4M instr/step): below the ~4% layout-noise measurement floor
That applies a cycles/binary-layout noise floor to an instruction-count measurement. The two floors are three orders of magnitude apart.
Measured on the C-LEARN run across six independent build+run pairs of identical source (Ryzen 9950X):
channel sd across builds retired instructions 0.026% branches 0.028% cycles, quiet machine 1.65% cycles, machine under load 9.9%–11% A null control makes the same point directly — the identical binary run as both sides of an interleaved A/B, 5 rounds, medians:
instructions -0.003% branches -0.004% cycles -1.540% <- an apparent "win" from nothingRetired instructions are a property of the program; cycles are a property of the machine executing it. So ~1.5% of instructions is roughly 58 sigma on that channel — comfortably measurable, not below any floor. The ~4% figure is real, but it bounds wall-clock/cycles claims from a single build pair, not instruction counts.
What the verdict actually rests on
Two reasons survive, and they are sufficient:
- Small absolute win for the highest design cost of the three candidates. Lazy
Ifneeds forward-jump opcodes, which touch codegen,max_stack_depthjoin validation, the peephole/fusion jump maps, the symbolic layer, and wasmgen parity. ~1.5% of instructions does not pay for that. - The cheap part of the win is already taken. Fusing
SetCond;If[;AssignCurr]into conditional-select opcodes removes 12.0% of executed dispatches against this issue's projected 15.9%, with none of that machinery — no jumps, no stack-depth join reasoning, no symbolic or wasmgen changes. It rests on the pair being adjacent by construction (compiler::codegen'sExpr::Ifarm is the sole producer of both and emits them together; measured executed counts are exactly equal at 1,874,169 each on C-LEARN).
So the remaining upside here is the residual after that fusion, against the full forward-jump design cost.
Why this is worth a comment rather than a silent doc edit
A wrong measurement premise in a design doc does not get caught by review: reviewers check the reasoning against the stated premise, not the premise against the world. "Below the measurement floor" reads as a fact and ends the discussion. This one had already carried a verdict, which is the evidence it was doing damage rather than sitting inert.
Anyone picking this issue up would have read the invalid floor claim here first, so fixing only the doc would have preserved the trap in the place most likely to be read.
The general convention that would have prevented it: when a doc or issue records a measurement-derived verdict, state which channel the number came from. A bare "only ~1.5%" invites comparison against whatever floor the reader has in mind.
- Small absolute win for the highest design cost of the three candidates. Lazy
- added a commit that references this issue
on Aug 11, 2026
Context
Identified during the simlin-engine VM performance round 2 (see
docs/design/engine-performance.md"Round 2 wins (2026-06-03)" and the unprioritized run-side swings at the bottom of that file). Sibling perf issues from the same campaign: #601 (threaded dispatch), #602 (lookup hot path), #603 (stride-step flat_offset), #604 (bundle marginal branch/dispatch reductions).Problem
The bytecode lowers
IF THEN ELSE(c, t, f)to "evaluatet; evaluatef; evaluatec;SetCond;If" — both branches are evaluated every timestep and one result is discarded. On C-LEARN there are 3,046Ifsites which, counting theirSetCondcondition chains, are roughly 10% of opcodes in the per-step flow program. Every step pays for the work of the branch that is not selected.Skipping the untaken branch is value-preserving: the discarded branch cannot affect the selected result (no observable side effects in the VM's expression evaluation), so this is a pure work reduction, not a semantics change.
Why it matters
Performance: ~10% of the per-step opcode budget on the largest model is spent computing values that are immediately thrown away. The run is throughput-bound on the round-2 machine (IPC ~4.5), so cutting executed work per step is the lever.
Scope / why it's a substantial design effort, not a quick patch
Lazy evaluation requires forward-jump opcodes, which touches several subsystems that today all assume straight-line code:
Expr::Ifarm insrc/simlin-engine/src/compiler/codegen.rsmust emit a conditional forward jump over the untaken branch instead of eagerly lowering both arms plusSetCond/If.max_stack_depthvalidation — the straight-line stack-depth computation inbytecode.rsassumes a single linear path; divergent stack depths across branches (the two arms push, then one is skipped) break that invariant and the validator needs to reason about join points.peephole_optimizeandfuse_three_addressmust not fuse across, or relocate, a jump target; they need jump-aware boundaries.CompiledSimulationstays symbolizable/disassemblable.SetCond/Ifeagerly too, so it needs the matching control-flow lowering (br_if/blocks) to stay behaviorally identical to the VM.Possible approach
Add forward-jump opcodes (conditional + unconditional), lower
Expr::Ifto "evalc; jump-if-false overt; evalt; jump overf; evalf". Extendmax_stack_depthto validate per-branch and at join points. Teach peephole/fusion and the symbolic layer about jump boundaries. Mirror the control flow in wasmgen. Gate the whole thing behind the existing simulate-suite cross-simulator tolerance plus VM-vs-wasm parity tests.Related context / cautionary note
A provably-equivalent rewrite of
vector_elm_map's strict-slice base as a precomputed affine dot product (structurally less work per element) measured a consistent ~5 ms REGRESSION on the 137 ms C-LEARN run — codegen perturbation of the giant inlinedeval_bytecodeswamped the algorithmic win. Documented indocs/design/engine-performance.mdas a negative result reinforcing #604's "measure everything near the eval loop" rule. Any change of this size to the eval loop's control flow must be measured, not argued: the structural win may not survive contact with the inliner.Refs
docs/design/engine-performance.md— "Round 2 wins (2026-06-03)" and the "LazyIf" bullet under the unprioritized run-side swings.src/simlin-engine/src/compiler/codegen.rs—Expr::Iflowering.src/simlin-engine/src/bytecode.rs—max_stack_depthvalidation,peephole_optimize,fuse_three_address.src/simlin-engine/src/vm.rs—SetCond/Ifhandlers.SetCond/Iflowering needing parity.