Repository navigation
ltm-finding: strongest-path discovery infeasible on large arrayed models (element-level blowup) #647
Description
Activity
- addedltmLoops that Matter (LTM) analysis subsystemLoops that Matter (LTM) analysis subsystem
on May 31, 2026 - added a commit that references this issue
on Jun 1, 2026 Update (2026-05-31, LTM-on-C-LEARN experience session): findings that raise the priority of this issue
1. The element-level blowup was not just slow -- it silently corrupted ALL results
The element-expanded layout doesn't just make discovery intractable; on C-LEARN it produced
garbagefor every LTM result anyone has ever observed on this model. C-LEARN + LTM-discovery instrumentation needs 171,498 result slots, but bytecodeVariableOffsetisu16(max 65,536 slots). Offsets past the limit silently wrapped (off as u16), so the variable at offset 65,536 overwrote slot 0 (time). The consequences:- Simulated time never advanced.
- Every saved row was identical.
- Every LTM result ever observed on C-LEARN -- link scores, loop scores, and the data in the previous experience report -- was garbage built on a non-advancing clock.
As of commit
aea4b9a5on branchltm-experience-and-improvements, this is a fast, clear compile error instead of silent corruption. But the corollary is that LTM is entirely unavailable on C-LEARN-scale models until the instrumentation fits under the u16 offset limit -- which makes the granularity fix proposed in this issue a correctness prerequisite, not just a performance optimization.2. Where the slots go, and two paths to shrink the count
The non-LTM C-LEARN layout is 5,215 slots; with discovery-mode LTM it is 171,498 slots (33x). The per-element link-score expansion is the multiplier. Two paths can shrink it dramatically:
- (a) Fix the MDL importer. The MDL importer expands a single A2A Vensim equation into N identical per-element equations (being filed as a separate issue). That expansion forces element-level rather than Bare-A2A reference sites in the LTM IR. Fixing the importer should collapse much of the element-graph back to variable level, attacking the slot count at the source.
- (b) Variable-level discovery -- this issue's existing proposed scope.
Either path reduces V; together they should get C-LEARN well under the u16 ceiling and into tractable discovery range.
3. Wall-clock breakdown (release build)
- LTM-discovery compile: 52.5 s
- VM run: 2.8 s
- (vs 0.6 s total without LTM)
The compile cost now dominates and is driven by generating and compiling the ~150k link-score equation fragments -- another consequence of the element-level expansion, and another reason the granularity reduction matters beyond DFS feasibility.
4. Discovery DFS budget is now enforced mid-step
Commit
7a68badmakes the discovery DFS budget enforced mid-step. Before that fix,Model.analyze(timeout=30)on C-LEARN burned 9+ CPU-minutes without honoring the timeout, because the budget was only checked between timesteps and a single step's DFS never finished. This bounds the worst case (the timeout is now respected), but it does not make discovery produce useful results at this scale -- it just fails fast instead of hanging. The underlying granularity/pruning redesign in this issue is still required.- added 5 commits that reference this issue
on Jun 1, 2026 Fixed by 081e984 ("engine: make LTM strongest-path discovery feasible on large models"). Three amendments (zero-score-edge exclusion, per-SCC DFS restriction, a component-scaled per-node expansion cap replacing the paper's best_score pruning) make element-level discovery tractable; see the "Scalability Amendments (GH #647)" section of docs/design/ltm--loops-that-matter.md.
Verified empirically on main (521fc37), release build: the full 251-step C-LEARN v77 discovery DFS completes in 0.07s and finds 153 loops (truncated: false), peak RSS ~390MB for the whole run — was >9.7 min for 4 steps / never completing for the full run.
- added a commit that references this issue
on Jun 2, 2026
Summary
LTM strongest-path loop discovery (
discover_loops_with_graphinsrc/simlin-engine/src/ltm_finding.rs) does not complete in feasible time on large arrayed models. Discovery is the heuristic path the engine auto-flips to when a model's causal-graph SCC exceedsMAX_LTM_SCC_NODES = 50(replacing exhaustive Johnson enumeration). On large arrayed models the discovery runs on the element-expanded causal graph, where the node count V reaches the tens of thousands, and it degenerates.This is distinct from #540 (World3): #540 is the scalar, variable-level dense SCC case (variable-level SCC of 166 scalar nodes). This issue is the element-level expansion blow-up on arrayed models, with C-LEARN as a now-working fixture, and a different proposed fix (run discovery at variable granularity for large arrayed models). The two share a root mechanism (the strongest-path pruning fails when path-score products do not shrink) but have distinct fixtures, measurements, and remediation. Tracked under epic #488; related to #540 and #481.
Measurements (C-LEARN v77,
test/xmutil_test_models/C-LEARN v77 for Vensim.mdl)The LTM sim of C-LEARN produces:
Per-timestep discovery cost explodes on "busy" timesteps:
This persists after a constant-factor optimization: the
IndexedSearchinteger-indexed-graph + reusable-scratch refactor of the per-timestep DFS (correct, keeps all LTM tests passing, cross-checked againstSearchGraphas an equivalence oracle). So the wall is algorithmic, not constant factors.Root-cause hypothesis (being tested separately)
score < best_score, which assumes path-score products shrink as paths extend. C-LEARN has link scores >> 1 (e.g. the ~5.2e7 degenerate macro-internal scores), so products can GROW, defeating the pruning and pushing toward near-exhaustive path enumeration.Proposed algorithmic redesign (the thing to track)
docs/reference/ltm--loops-that-matter.mdsection 12.4) also notes discovery can be run at a subset of timesteps -- another lever.Why it matters
Discovery is meant to be the tractable path for large models, but on a real arrayed model (C-LEARN) it does not finish at all. Any large arrayed model that auto-flips to discovery (SCC >
MAX_LTM_SCC_NODES) will hang in the production analysis path. Correctness of the per-timestep DFS is fine; the issue is purely scalability/feasibility at element-expanded scale.Components affected
src/simlin-engine/src/ltm_finding.rs(discover_loops_with_graph, the per-timestep strongest-path DFS,IndexedSearch)docs/reference/ltm--loops-that-matter.mdsection 12 (algorithm reference)src/simlin-engine/examples/clearn_discover.rs(env varsCLEARN_SKIP_DISCOVERY,CLEARN_DISCOVER_STEPS=N,CLEARN_CAP_SCORES)How it was discovered
Diagnosed while profiling LTM discovery on C-LEARN v77 using the
clearn_discover.rsphase-timing harness, after confirming theIndexedSearchconstant-factor refactor did not move the wall.References
src/simlin-engine/src/ltm_finding.rsdocs/reference/ltm--loops-that-matter.mdsection 12src/simlin-engine/examples/clearn_discover.rs