You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
ltm: --ltm warning order is process-random (model_ltm_variables iterates CausalEdgesResult.edges HashMap) #1036
simlin simulate --ltm prints its LTM warnings to stderr in a different order on every process, for the same binary and the same model. The warning set is stable; only the order moves.
Reproduced with a release build at c8770abb (which has #1003 / d5db88a3 as an ancestor) on test/xmutil_test_models/C-LEARN v77 for Vensim.mdl, three fresh processes:
all three runs emit exactly 25 stderr lines;
the sorted files are byte-identical (cmp of sorted output: equal for 1==2 and 1==3);
the raw files differ pairwise (1!=2, 1!=3, 2!=3), first divergence at line 4 in every pair.
The shuffled lines are the model_ltm_variables warnings: the GH #758 "LTM link score for edge X -> Y could not be computed: both variables are arrayed but their dimensions do not correspond" rows and the GH #995RankLikePartial "LTM link-score variable '...' could not be generated: ... applies an array-producing builtin" rows. The lines above them (the # Loops in model listing and the auto-flip-to-discovery advisory) keep their position.
Root cause (verified by reading the code, and consistent with the observed shape)
model_ltm_variables (src/simlin-engine/src/db/ltm/mod.rs:1516) drives link-score emission for discovery-mode and input-port models with
for(from, tos)in&edges_result.edges{for to in tos {emit_link_scores_for_edge(db, ..., from, to, ...,&mut vars,&mut unscoreable_edges);}}
and CausalEdgesResult.edges is HashMap<String, BTreeSet<String>> (src/simlin-engine/src/db/analysis.rs:39). Every emit_unscoreable_*_warning / rank-like decline in db/ltm/link_scores.rs accumulates its CompilationDiagnostic from inside that loop, the salsa accumulator drain preserves emission order, collect_model_diagnostics (db/diagnostic.rs:1179) dedups but does not sort, and the CLI prints the Vec as-is. So the warning order is the HashMap's per-process RandomState order.
The observation matches the type exactly: warnings sharing a from variable (population -> aggregated_population, population -> semi_agg_population) are adjacent in every run (inner BTreeSet is ordered) while the from groups shuffle (outer HashMap is not).
C-LEARN takes this branch because it auto-flips to discovery mode (its largest variable-level SCC has 149 nodes, above MAX_LTM_SCC_NODES = 50). Exhaustive-mode root models take the other branch (for loop_item in detected_loops), which is ordered.
Two more hash-ordered loops on the same path, not exercised by the C-LEARN root model but the same defect:
db/ltm/mod.rs:2030: for (input_port, port_pathways) in &pathways -- pathways is a HashMap (built at mod.rs:1098 / module_input_pathways_from_edges), so the unresolved-pathway-edge warnings of a model with input ports are also process-ordered.
src/simlin-cli/src/main.rs:522: for (model_name, source_model) in models.iter() over project.models(db) (a HashMap, per the comment on the collect_all_diagnostics sort) prints the # Loops in model '...' blocks in process order on a multi-model project.
The synthetic-var list itself is fine: vars.sort_by at mod.rs:2160 orders it, so model_ltm_fragment_diagnostics (which iterates ltm_vars.vars) and its name-sorted implicit-helper leg emit deterministically.
#999 was closed by #1003 with "every HashMap-iteration order that reached the diagnostics collection is sorted", verified by a 931-row corpus check that was byte-identical across runs. That sweep ran collect_all_diagnostics with ltm_enabled = false (the default), and none of the four fixtures in src/simlin-engine/src/db/diagnostic_determinism_tests.rs enable LTM (the file has zero mentions of ltm). The model_ltm_variables warnings are emitted only inside the project.ltm_enabled(db) gate, so they were outside both the sweep and the fixtures. This is the residual of the #999 class on the LTM-only path, not a regression of the #1003 fix.
Why it matters
Low severity: text only, no numbers move, and the set is stable. But:
any "diagnostics are stable across an edit / across builds" gate on an --ltm run flaps (this is how it was noticed: diffing --ltm stderr between two builds of the same source required sorting first);
Sort once in the collector: order collect_model_diagnostics's output by (model, variable, rendered message). Closes every current and future site at once, but Diagnostic derives only PartialEq, Eq, Hash (no Ord), so it needs an explicit sort key, and it changes the grouping of the non-LTM passes engine: LTM array-freeze materialization, diagnostics attribution/determinism, and group-aware ranking #1003 deliberately kept in execution order.
Make CausalEdgesResult.edges a BTreeMap / IndexMap. Broadest, but edges is on the LTM hot path and is consumed in many places; measure before choosing this.
Whichever approach: add an LTM-enabled fixture to diagnostic_determinism_tests.rs (set_project_ltm_enabled, a discovery-mode or input-port model with at least two unscoreable edges from different from variables, using the file's fresh-database repetition recipe) and mutation-check it by reverting the sort, as #1003 did for its fixtures.
Discovery context
Observed during the compiler-unification Phase 7.4 review on branch compiler-unification-v2, where two builds of the same source printed the same 25 warnings in different orders. Pre-existing: reproduced above from three processes of one binary, with the #1003 fix present in history.
Summary
simlin simulate --ltmprints its LTM warnings to stderr in a different order on every process, for the same binary and the same model. The warning set is stable; only the order moves.Reproduced with a release build at
c8770abb(which has #1003 /d5db88a3as an ancestor) ontest/xmutil_test_models/C-LEARN v77 for Vensim.mdl, three fresh processes:cmpofsorted output: equal for 1==2 and 1==3);The shuffled lines are the
model_ltm_variableswarnings: the GH #758 "LTM link score for edge X -> Y could not be computed: both variables are arrayed but their dimensions do not correspond" rows and the GH #995RankLikePartial"LTM link-score variable '...' could not be generated: ... applies an array-producing builtin" rows. The lines above them (the# Loops in modellisting and the auto-flip-to-discovery advisory) keep their position.Root cause (verified by reading the code, and consistent with the observed shape)
model_ltm_variables(src/simlin-engine/src/db/ltm/mod.rs:1516) drives link-score emission for discovery-mode and input-port models withand
CausalEdgesResult.edgesisHashMap<String, BTreeSet<String>>(src/simlin-engine/src/db/analysis.rs:39). Everyemit_unscoreable_*_warning/ rank-like decline indb/ltm/link_scores.rsaccumulates itsCompilationDiagnosticfrom inside that loop, the salsa accumulator drain preserves emission order,collect_model_diagnostics(db/diagnostic.rs:1179) dedups but does not sort, and the CLI prints theVecas-is. So the warning order is theHashMap's per-processRandomStateorder.The observation matches the type exactly: warnings sharing a
fromvariable (population -> aggregated_population,population -> semi_agg_population) are adjacent in every run (innerBTreeSetis ordered) while thefromgroups shuffle (outerHashMapis not).C-LEARN takes this branch because it auto-flips to discovery mode (its largest variable-level SCC has 149 nodes, above
MAX_LTM_SCC_NODES = 50). Exhaustive-mode root models take the other branch (for loop_item in detected_loops), which is ordered.Two more hash-ordered loops on the same path, not exercised by the C-LEARN root model but the same defect:
db/ltm/mod.rs:2030:for (input_port, port_pathways) in &pathways--pathwaysis aHashMap(built atmod.rs:1098/module_input_pathways_from_edges), so the unresolved-pathway-edge warnings of a model with input ports are also process-ordered.src/simlin-cli/src/main.rs:522:for (model_name, source_model) in models.iter()overproject.models(db)(aHashMap, per the comment on thecollect_all_diagnosticssort) prints the# Loops in model '...'blocks in process order on a multi-model project.The synthetic-var list itself is fine:
vars.sort_byatmod.rs:2160orders it, somodel_ltm_fragment_diagnostics(which iteratesltm_vars.vars) and its name-sorted implicit-helper leg emit deterministically.Why #999 did not cover this
#999 was closed by #1003 with "every HashMap-iteration order that reached the diagnostics collection is sorted", verified by a 931-row corpus check that was byte-identical across runs. That sweep ran
collect_all_diagnosticswithltm_enabled = false(the default), and none of the four fixtures insrc/simlin-engine/src/db/diagnostic_determinism_tests.rsenable LTM (the file has zero mentions ofltm). Themodel_ltm_variableswarnings are emitted only inside theproject.ltm_enabled(db)gate, so they were outside both the sweep and the fixtures. This is the residual of the #999 class on the LTM-only path, not a regression of the #1003 fix.Why it matters
Low severity: text only, no numbers move, and the set is stable. But:
--ltmrun flaps (this is how it was noticed: diffing--ltmstderr between two builds of the same source required sorting first);Components
src/simlin-engine/src/db/ltm/mod.rs(model_ltm_variables, lines 1516 and 2030)src/simlin-engine/src/db/analysis.rs(CausalEdgesResult.edgestype)src/simlin-cli/src/main.rs(simulate(.., enable_ltm)loop listing, line 522)Possible approaches
edges_result.edgeskeys into aVec<&String>, sort, iterate; same forpathways; sortmodelsin the CLI (or reuse the sorted-models iterationcollect_all_diagnosticsalready does). Smallest change, keeps the per-site "query-execution order" grouping engine: LTM array-freeze materialization, diagnostics attribution/determinism, and group-aware ranking #1003 chose. Cost is one sort of thefromset permodel_ltm_variablesevaluation, negligible next to the compile itself.collect_model_diagnostics's output by(model, variable, rendered message). Closes every current and future site at once, butDiagnosticderives onlyPartialEq, Eq, Hash(noOrd), so it needs an explicit sort key, and it changes the grouping of the non-LTM passes engine: LTM array-freeze materialization, diagnostics attribution/determinism, and group-aware ranking #1003 deliberately kept in execution order.CausalEdgesResult.edgesaBTreeMap/IndexMap. Broadest, butedgesis on the LTM hot path and is consumed in many places; measure before choosing this.Whichever approach: add an LTM-enabled fixture to
diagnostic_determinism_tests.rs(set_project_ltm_enabled, a discovery-mode or input-port model with at least two unscoreable edges from differentfromvariables, using the file's fresh-database repetition recipe) and mutation-check it by reverting the sort, as #1003 did for its fixtures.Discovery context
Observed during the compiler-unification Phase 7.4 review on branch
compiler-unification-v2, where two builds of the same source printed the same 25 warnings in different orders. Pre-existing: reproduced above from three processes of one binary, with the #1003 fix present in history.