Skip to content
45 changes: 45 additions & 0 deletions source/compiler/qsc/src/codegen/tests.rs
Original file line number Diff line number Diff line change
Expand Up @@ -91,6 +91,51 @@ fn compile_source_to_qir_result(
)
}

#[test]
fn recursive_callable_factory_is_inlined_in_adaptive_qir() {
let source = r#"
namespace Test {
operation Flip(q : Qubit) : Unit { X(q); }

function Make(n : Int) : Qubit => Unit {
if n == 0 {
Flip
} else {
let next = Make(n - 1);
next
}
}

@EntryPoint()
operation Main() : Result {
use q = Qubit();
Make(2)(q);
MResetZ(q)
}
}
"#;
for profile in [Profile::Adaptive, Profile::AdaptiveRIF] {
let capabilities = profile.into();
let mut interpreter = interpreter_with_capabilities(capabilities);
eval_fragments(&mut interpreter, source);
let incremental_qir = interpreter
.qirgen("Test.Main()")
.unwrap_or_else(|errors| panic!("{}", format_interpret_errors(errors)));
for qir in [compile_source_to_qir(source, capabilities), incremental_qir] {
assert!(!qir.contains("@Make"), "factory must be inlined: {qir}");
assert_eq!(
qir.matches("call void @__quantum__qis__x__body").count(),
1,
"{qir}"
);
assert!(
qir.contains("call void @__quantum__rt__result_record_output"),
"{qir}"
);
}
}
}

#[test]
fn generic_lambda_dependencies_compile_through_executable_entries() {
check_generic_lambda_dependency_qir(compile_source_to_qir);
Expand Down
53 changes: 41 additions & 12 deletions source/compiler/qsc_fir_transforms/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,32 +6,54 @@ The production FIR-to-FIR rewrite pipeline. It runs after FIR lowering and befor

- **It is one pipeline, not a toolbox of independent passes.** The passes are ordered and assume each other's output. Several intermediate states deliberately violate FIR invariants that later passes restore, so running a pass in isolation or reordering passes is generally unsound. Treat `run_pipeline_with_diagnostics` (and the staged `run_pipeline_to_with_diagnostics`) as the only supported way to invoke them.

- **Rewrites are entry-reachability-driven.** Most passes inspect what is reachable from the package entry expression and only mutate that. UDT erasure is the main exception: it is still reachability-scoped but works at package granularity across the reachable package closure (target package plus any package with an entry-reachable callable; unreachable packages are left alone).
- **Rewrites are reachability-driven.** Most passes start from the package entry expression. UDT erasure includes pinned roots and works at package granularity across that reachable closure. Argument promotion includes pinned bodies and dependencies in safety checks and caller rewrites, but preserves pinned root signatures.

- **One `Assigner` is threaded through the whole pipeline.** Every pass that synthesizes FIR nodes allocates fresh IDs from a single shared `Assigner` so IDs never collide across stages. Do not construct a new `Assigner` mid-pipeline. The trailing metadata passes (`gc_unreachable`, `item_dce`, `exec_graph_rebuild`) don't get it because they only tombstone, delete, or rebuild derived data.
- **One `PackageAssigners` pool is threaded through the pipeline.** Each package has its own ID space and reuses its own `Assigner` across passes. Allocate copied guards and other synthesized nodes with the owning package's assigner, not the entry package's. The trailing metadata passes only delete nodes or rebuild derived data.

- **Synthesized nodes use the `EMPTY_EXEC_RANGE` sentinel.** New exprs/stmts carry an empty `exec_graph_range`; the final `exec_graph_rebuild` pass consumes that sentinel and recomputes the execution graph.

- **Only consume output when there are no fatal diagnostics.** Fatal diagnostics (from `return_unify`, `defunctionalize`, or pinned-item validation) leave the store at an intermediate, invalid state. Warning-only diagnostics are preserved and do not block successful output.
- **Only consume output when there are no fatal diagnostics.** A failed pipeline can leave the store at an intermediate, invalid state. Warning-only diagnostics are preserved and do not block successful output.

## Pass order

The driver first validates intrinsics, collapses simulatable intrinsics, and clears orphaned nodes. The main rewrite schedule is:

1. `monomorphize` — specialize reachable generic callables to concrete types.
2. `return_unify` — rewrite bodies to single-exit form, removing `Return` nodes while preserving path-local side effects (e.g. qubit release).
3. `defunctionalize` — eliminate callable-valued expressions/closures; rewrite call sites to direct dispatch.
4. `udt_erase` — replace UDT values and struct expressions with tuple/scalar form across the reachable package closure.
5. `tuple_compare_lower` — lower equality/inequality on non-empty tuples to element-wise scalar comparisons.
6. `tuple_decompose` — decompose tuple-valued locals whose uses are all field accesses.
7. `arg_promote` — flatten tuple-valued callable parameters and update call sites.
3. `cond_normalize` — preserve selection-time conditions before callable analysis and dispatch rewriting.
4. `defunctionalize` — specialize known callable choices and rewrite calls to direct dispatch. Unresolved alternatives remain dynamic rather than being discarded in favor of a known branch.
5. `udt_erase` — replace UDT values and struct expressions with tuple/scalar form across the reachable package closure. Core Complex addition, subtraction, multiplication and unary signs lower to scalar component operations before nominal identity is lost.
6. `tuple_compare_lower` — lower equality/inequality on non-empty tuples to element-wise scalar comparisons.
7. `tuple_decompose` — decompose eligible tuple-valued locals into scalar fields.
8. `arg_promote` — flatten tuple-valued callable parameters and update call sites.

Steps 6 and 7 iterate to a fixed point (convergence is guaranteed by a strictly-decreasing measure).
Steps 7 and 8 iterate to a fixed point (convergence is guaranteed by a strictly-decreasing measure).

8. `gc_unreachable` — tombstone orphaned arena nodes.
9. `item_dce` — remove unreachable callable/type items; re-run `gc_unreachable` if anything was deleted.
10. `exec_graph_rebuild` — recompute exec-graph metadata from the rewritten FIR.
9. `normalize_reachable_call_arg_types` — reconcile call argument types once after the fixed point.
10. `item_dce` — remove unreachable items while retaining pinned callables and their dependencies.
11. `gc_unreachable` — remove orphaned arena nodes across the reachable package closure.
12. `exec_graph_rebuild` — recompute exec-graph metadata from the rewritten FIR.

Invariant checks run after most passes. `run_pipeline_to_with_diagnostics` exposes each stage as a cut point used by tests and (with `PipelineStage::Full` plus pinned callable items) by production codegen.

Within defunctionalization, higher-order and direct-call rewrites share a
children-before-parents traversal. An enclosing call must copy already-rewritten
arguments, not stale argument layouts paired with a newly specialized callee.
If an inner rewrite replaces a capture operand with a new local read, its
dependent direct or higher-order calls wait for fresh capture analysis. Direct calls also wait
when their callee gains control flow that needs normalization. Closure cleanup treats
computed callees as live dependencies, just like call arguments; consuming one
use does not make another invocation disposable.

Complex component lowering currently covers addition, subtraction, multiplication,
unary signs, and the corresponding compound assignments. Division and
exponentiation are not covered by this lowering.

Callable capability weakening does not change value layout: an adjointable or
controllable operation can populate a less-capable operation binding without
adding tuple wrappers. Assignment checks retain nominal identity, tuple shape,
callable kind and input/output compatibility; they do not allow capability upgrades.

## Where to look

- `src/lib.rs` — pipeline orchestration, stage cut points, and the cross-pass contracts above.
Expand All @@ -49,3 +71,10 @@ cargo test -p qsc_fir_transforms --features slow-proptest-tests # + semantic-e
```

Pass-local unit tests sit next to each pass; `tests/pipeline_integration.rs` drives full-pipeline and per-stage behavior.

Return-normalization semantic tests compare explicit expected values, ordered
quantum-operation and lifetime traces, and receiver output before and after the
Full pipeline. They do not establish QIR support by themselves. The
short-circuit-assignment codegen tests keep qubit ownership in the caller so they
exercise return lowering independently of conditional qubit cleanup. The
resource-owning #3836 case still requires the separate cleanup repair.
50 changes: 38 additions & 12 deletions source/compiler/qsc_fir_transforms/src/arg_promote.rs
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,9 @@
//! `Var(Res::Item)` with `Ty::Arrow` outside a `Call` callee position) or as
//! a closure target, since indirect dispatch requires a stable parameter
//! layout (this also covers partial-application cases).
//! - **Pinned callers:** entry and pinned roots share safety analysis and caller
//! rewriting. Pinned roots keep their signatures; their callees may promote
//! only when every retained use can follow the changed layout.
//! - **Per iteration:** reachability scan → eligibility analysis
//! ([`check_candidates`]) → safety filters
//! ([`collect_first_class_callables`], [`collect_closure_targets`]) →
Expand Down Expand Up @@ -55,7 +58,7 @@ use crate::fir_builder::{
functored_specs,
};
use crate::package_assigners::PackageAssigners;
use crate::reachability::collect_reachable_from_entry;
use crate::reachability::collect_reachable_with_seeds;
use crate::walk_utils::{
ParamUse, classify_uses_in_block, collect_expr_ids_in_entry_and_local_callables,
collect_expr_ids_in_local_callables, for_each_expr, for_each_expr_in_callable_impl,
Expand Down Expand Up @@ -132,20 +135,32 @@ type ParamLeafRemap = (LocalVarId, Ty, LeafRemap);
/// # Panics
///
/// Panics if the package has no entry expression. The reachability scans
/// in this pass go through [`collect_reachable_from_entry`], which asserts
/// in this pass go through [`collect_reachable_with_seeds`], which asserts
/// `package.entry.is_some()`.
#[cfg(test)]
pub fn arg_promote(
store: &mut PackageStore,
package_id: PackageId,
assigners: &mut PackageAssigners,
) -> bool {
arg_promote_with_pins(store, package_id, assigners, &[])
}

/// Includes pinned bodies and dependencies in safety checks and caller rewrites,
/// while retaining the pinned roots' externally visible input signatures.
pub(crate) fn arg_promote_with_pins(
store: &mut PackageStore,
package_id: PackageId,
assigners: &mut PackageAssigners,
pinned: &[StoreItemId],
) -> bool {
let mut tmp_counter: u32 = 0;
let changed = promote_to_fixed_point(store, package_id, assigners, &mut tmp_counter);
normalize_reachable_call_arg_types(store, package_id, assigners);
let changed = promote_to_fixed_point(store, package_id, assigners, &mut tmp_counter, pinned);
normalize_reachable_call_arg_types(store, package_id, assigners, pinned);
changed
}

/// Iterates promotion rounds until no more candidates are found.
/// Iterates entry- and pin-reachable promotion rounds until no candidates remain.
///
/// Each iteration peels one level of tuple nesting from eligible parameters,
/// rewrites their bodies and call sites, then recomputes reachability for
Expand All @@ -163,26 +178,35 @@ pub(crate) fn promote_to_fixed_point(
package_id: PackageId,
assigners: &mut PackageAssigners,
tmp_counter: &mut u32,
pinned: &[StoreItemId],
) -> bool {
let mut changed = false;
loop {
let candidates = find_promotion_candidates(store, package_id);
let candidates = find_promotion_candidates(store, package_id, pinned);
if candidates.is_empty() {
break;
}
changed = true;
apply_promotions(store, package_id, assigners, &candidates, tmp_counter);
apply_promotions(
store,
package_id,
assigners,
&candidates,
tmp_counter,
pinned,
);
}
changed
}

/// Finds all eligible promotion candidates in the current reachable set,
/// excluding callables used as first-class values or closure targets.
/// excluding pinned roots, first-class values and closure targets.
fn find_promotion_candidates(
store: &PackageStore,
package_id: PackageId,
pinned: &[StoreItemId],
) -> Vec<ArgPromoCandidate> {
let reachable = collect_reachable_from_entry(store, package_id);
let reachable = collect_reachable_with_seeds(store, package_id, pinned);

// The entry callable lives in the true entry package only; resolving it
// there keeps its input ABI excluded from flattening regardless of how many
Expand Down Expand Up @@ -211,7 +235,7 @@ fn find_promotion_candidates(
// The entry-point callable's input is the program's externally-visible
// ABI and must never be flattened, regardless of its input shape.
// This is a forward looking check as all inputs are currently `Unit`
if Some(owner) == entry_item {
if Some(owner) == entry_item || pinned.contains(&owner) {
continue;
}
if first_class.contains(&owner) || closure_targets.contains(&owner) {
Expand Down Expand Up @@ -246,6 +270,7 @@ fn apply_promotions(
assigners: &mut PackageAssigners,
candidates: &[ArgPromoCandidate],
tmp_counter: &mut u32,
pinned: &[StoreItemId],
) {
// Group candidates by their declaring callable so each callable's entire
// input is flattened exactly once, dissolving all inter-parameter
Expand Down Expand Up @@ -294,7 +319,7 @@ fn apply_promotions(
let promoted_map: FxHashMap<StoreItemId, PromotionResult> =
promotions.into_iter().map(|p| (p.item_id, p)).collect();

let reachable = collect_reachable_from_entry(store, package_id);
let reachable = collect_reachable_with_seeds(store, package_id, pinned);
let mut caller_pkgs: Vec<PackageId> = Vec::new();
for store_id in &reachable {
if !caller_pkgs.contains(&store_id.package) {
Expand Down Expand Up @@ -336,8 +361,9 @@ pub(crate) fn normalize_reachable_call_arg_types(
store: &mut PackageStore,
package_id: PackageId,
assigners: &mut PackageAssigners,
pinned: &[StoreItemId],
) {
let reachable = collect_reachable_from_entry(store, package_id);
let reachable = collect_reachable_with_seeds(store, package_id, pinned);

// Snapshot every reachable callable's current input type, package-qualified,
// so a direct call to a foreign promoted callee resolves to the callee's own
Expand Down
47 changes: 22 additions & 25 deletions source/compiler/qsc_fir_transforms/src/arg_promote/tests.rs
Original file line number Diff line number Diff line change
Expand Up @@ -3080,21 +3080,19 @@ fn build_leaf_tuple_interior_whole_tuple_read_preserves_values() {
}

#[test]
fn whole_tuple_copy_assignment_is_decomposed() {
// `set x = y;` copies the whole tuple value `y` into `x`. By `TupleDecompose2`
// the copy-assignment normalization and tuple-decompose have split the copy into
// per-element assignments and scalar-replaced both `x` and `y`, leaving no
// `(Int, Int)` tuple local. The value semantics are guarded by the paired
// `whole_tuple_copy_assignment_preserves_evaluated_values` test below.
fn whole_tuple_copy_saves_source_fields_before_scalar_stores() {
// This pins the lowered shape; the paired semantic test below checks 34.
check_at_stage(
"function Main() : Unit { mutable x = (1, 2); let y = (3, 4); x = y; }",
PipelineStage::TupleDecompose2,
&expect![[r#"
function Main() : Unit {
mutable (x_0 : Int, x_1 : Int) = (1, 2);
let (y_0 : Int, y_1 : Int) = (3, 4);
x_0 = y_0;
x_1 = y_1;
let __tuple_rhs_0 : Int = y_0;
let __tuple_rhs_1 : Int = y_1;
x_0 = __tuple_rhs_0;
x_1 = __tuple_rhs_1;
}
// entry
Main()
Expand All @@ -3108,7 +3106,7 @@ fn whole_tuple_copy_assignment_preserves_evaluated_values() {
// place-valued elements make every position observable. With y = (3, 4) the
// copy `set x = y;` must yield x.0 = 3, x.1 = 4, so the result is
// 3 * 10 + 4 = 34. A swapped-index copy bug would change the number.
check_semantic_equivalence(
crate::test_utils::check_semantic_equivalence_with_expected(
"@EntryPoint()
function Main() : Int {
mutable x = (0, 0);
Expand All @@ -3117,6 +3115,7 @@ fn whole_tuple_copy_assignment_preserves_evaluated_values() {
let (a, b) = x;
a * 10 + b
}",
qsc_eval::val::Value::Int(34),
);
}

Expand Down Expand Up @@ -3144,28 +3143,25 @@ fn whole_tuple_copy_assignment_partial_decompose_with_whole_use() {
}

#[test]
fn nested_whole_tuple_copy_assignment_preserves_values() {
// A nested copy `set x = y;` where both are `(Int, (Int, Int))` fully
// decomposes to scalar leaves across the fixed point. The interesting part
// is that the inner copy is *regenerated* mid-loop: round 1 normalizes the
// top level to `set x = (y::0, y::1)` and tuple-decompose splits it into
// `set x_0 = y::0; set x_1 = y::1` while scalar-replacing `y`, which rewrites
// `y::1` into the bare `Var(y_1)`. That leaves a fresh `set x_1 = y_1`
// whole-tuple Var-to-Var copy, which the *next* round re-normalizes and
// decomposes. This is why copy-assignment normalization must run every
// fixed-point iteration rather than once up front. The end state shown here is
// stable by `TupleDecompose2`; the value semantics are guarded by the paired
// `nested_whole_tuple_copy_assignment_preserves_evaluated_values` test below.
fn nested_tuple_copy_scalarizes_rhs_snapshots_through_fixpoint() {
// Saving the inner tuple introduces another whole-value copy. Later rounds
// must normalize and scalarize that copy too. This asserts the final shape;
// the paired semantic test below checks 789.
check_at_stage(
"function Main() : Unit { mutable x = (0, (0, 0)); let y = (7, (8, 9)); x = y; }",
PipelineStage::TupleDecompose2,
&expect![[r#"
function Main() : Unit {
mutable (x_0 : Int, (x_1_0 : Int, x_1_1 : Int)) = (0, (0, 0));
let (y_0 : Int, (y_1_0 : Int, y_1_1 : Int)) = (7, (8, 9));
x_0 = y_0;
x_1_0 = y_1_0;
x_1_1 = y_1_1;
let __tuple_rhs_0 : Int = y_0;
let __tuple_rhs_1_0 : Int = y_1_0;
let __tuple_rhs_1_1 : Int = y_1_1;
x_0 = __tuple_rhs_0;
let __tuple_rhs_0_1 : Int = __tuple_rhs_1_0;
let __tuple_rhs_1 : Int = __tuple_rhs_1_1;
x_1_0 = __tuple_rhs_0_1;
x_1_1 = __tuple_rhs_1;
}
// entry
Main()
Expand All @@ -3178,7 +3174,7 @@ fn nested_whole_tuple_copy_assignment_preserves_evaluated_values() {
// Value guard for the nested copy: with y = (7, (8, 9)) the element-wise
// copy must yield 7 * 100 + 8 * 10 + 9 = 789. A cross-wired leaf copy would
// change the number.
check_semantic_equivalence(
crate::test_utils::check_semantic_equivalence_with_expected(
"@EntryPoint()
function Main() : Int {
mutable x = (0, (0, 0));
Expand All @@ -3187,5 +3183,6 @@ fn nested_whole_tuple_copy_assignment_preserves_evaluated_values() {
let (a, (b, c)) = x;
a * 100 + b * 10 + c
}",
qsc_eval::val::Value::Int(789),
);
}
Loading