You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Running list: potential discrepancies in preprint timings #183
Follow-up work on the differentiable-solver benchmarks (dolfinx-adjoint#81, #179, #182) has surfaced several timing discrepancies between Table 3 of the paper (arXiv:2606.27895) and what the current benchmark suite reproduces. This issue collects them in one place and tracks their resolution.
These look like artifacts of the earlier benchmark harness and the environment the paper's numbers were produced in, not fundamental problems with the solvers or the physics. The forward and gradient accuracy results are unaffected. Only the wall-clock cost columns are in question. We intend to re-run the affected timing measurements on the current suite and hardware, incorporating the first wave of third-party solver fixes, and update the table accordingly.
The paper is still at the preprint stage, so correcting these now is straightforward. To be clear, this is a one-time reconciliation while the work is pre-publication, not an open-ended commitment to chase every measurement indefinitely.
Discrepancies
1. FEniCS structural forward time (~140 s) does not reproduce
Table 3 lists the legacy-FEniCS structural forward at ~140 s (N=128). On the current suite (#179 extended the cost sweep to the paper's 128×2×64 resolution), FEniCS forward is ~5.5 s and VJP ~16.3 s on CI. After #182 (UMFPACK→mumps plus tape caching) it drops to ~3.7 s / ~11.1 s. That is a ~25–38× gap from the published number.
The FEniCS solver code is essentially unchanged since it was written, so this is not a solver regression that was later fixed.
The most likely cause is the paper-repro environment (thread count, solver defaults, no warm-tape reuse) rather than the solver. An end-to-end re-run is needed to confirm.
2. Sub-1.0 VJP/forward ratios are not physically reachable
Several Table 3 rows report VJP cheaper than a standalone forward:
row
Fwd
VJP
ratio
FEniCS (H)
60 s
47 s
0.78
FEniCS (S)
140 s
86 s
0.61
torch-fem (H)
33 s
11 s
0.33
TopOpt.jl (S)
78 s
15 s
0.19
The VJP column times jax.grad, which dispatches one apply RPC followed by one vector_jacobian_product RPC, and most VJP endpoints re-run the forward internally. A VJP measurement is therefore about two forward solves plus one adjoint, which cannot come out below 1× a forward. JAX-FEM and Firedrake sit where this predicts (~2.0–2.4). The sub-1.0 rows should not be reachable. (Credit: @jorgensd, #182 §3.)
The current suite confirms this directly. Structural cost at nx=128 (paper resolution) on our CI shows every solver's VJP/forward ratio at 1.7 or above, including the ones the paper puts below 1.0:
solver (structural, nx=128)
fwd
vjp
vjp/fwd
Table 3 ratio
FEniCS
5.47 s
16.31 s
2.98
0.61
Firedrake
5.07 s
12.64 s
2.49
2.4
JAX-FEM
13.51 s
38.48 s
2.85
2.3
TopOpt.jl
0.63 s
1.06 s
1.69
0.19
Firedrake and JAX-FEM match the paper. FEniCS and TopOpt.jl flip from sub-1.0 to above 1.0 under re-measurement.
3. torch-fem row: re-measured at ~1.6×, not 0.33
We reproduced the torch-fem thermal endpoints in a container (torch-fem 0.6.2, CPU, which is the tesseract's actual numeric path since it hardcodes _SPARSE_SOLVE_DEVICE="cpu"), instrumenting every sparse_solve:
N
DOFs
apply
VJP
ratio
solves/apply
solves/VJP
64
4,290
~0.4 s
~0.6–1.1 s
1.4–2.5
1
2
128
16,770
~2.4–3.5 s
~3.4–4.8 s
~1.4
2
3
256
66,306
~7–9 s
~12–14.5 s
~1.60
2
3
VJP/forward is ~1.6 at the paper's N=256, not 0.33. The VJP always does one extra adjoint solve on top of the forward, exactly as the source implies. This is consistent with @jorgensd's in-process ~1.5.
The preconditioner reuse the paper describes is real. The adjoint reuses the AMG hierarchy, making it ~0.77× a forward. That reduces the adjoint cost but cannot invert the ratio below 1.
Our absolute forward at N=256 is single-digit seconds against the published 33 s. Like the FEniCS 140 s, that large gap points to the earlier environment, so the torch-fem row is worth re-running end-to-end rather than just checking its ratio.
Caveat on our numbers: CPU, in-process (not over the tesseract RPC), on a laptop, so absolute times carry run-to-run variance. The ratio stays above 1 at every size across repeats.
4. TopOpt.jl (0.19) — not yet checked
The listed strategy is "Analytic." For self-adjoint compliance minimisation the λ=−u shortcut means no separate adjoint solve, so a genuinely cheap VJP is plausible. It still needs a forward solution, so the same "cannot be far below a forward" question applies. Not yet independently measured.
changed the title [-]Reconcile Table 3 timing discrepancies with the current benchmark suite[/-][+]Running list: potential discrepancies in preprint timings[/+]on Sep 18, 2026
The statement direct (spsolve) in Table 2 is not strictly correct. On GPU, the thermal example uses CG + Jacobi (more iterations are faster than spending time on an advanced preconditioner here). On CPU, the thermal example dispatches based on DOF count: Below 10_000, it's direct (spsolve), above it's CG + AMG. That's why there was a jump in the table and the paper's N=256 was never a direct solve. I suggest the GPU CG-Jacobi path with a clear indication of GPU usage.
Summary
Follow-up work on the differentiable-solver benchmarks (dolfinx-adjoint#81, #179, #182) has surfaced several timing discrepancies between Table 3 of the paper (arXiv:2606.27895) and what the current benchmark suite reproduces. This issue collects them in one place and tracks their resolution.
These look like artifacts of the earlier benchmark harness and the environment the paper's numbers were produced in, not fundamental problems with the solvers or the physics. The forward and gradient accuracy results are unaffected. Only the wall-clock cost columns are in question. We intend to re-run the affected timing measurements on the current suite and hardware, incorporating the first wave of third-party solver fixes, and update the table accordingly.
The paper is still at the preprint stage, so correcting these now is straightforward. To be clear, this is a one-time reconciliation while the work is pre-publication, not an open-ended commitment to chase every measurement indefinitely.
Discrepancies
1. FEniCS structural forward time (~140 s) does not reproduce
Table 3 lists the legacy-FEniCS structural forward at ~140 s (N=128). On the current suite (#179 extended the cost sweep to the paper's 128×2×64 resolution), FEniCS forward is ~5.5 s and VJP ~16.3 s on CI. After #182 (UMFPACK→mumps plus tape caching) it drops to ~3.7 s / ~11.1 s. That is a ~25–38× gap from the published number.
2. Sub-1.0 VJP/forward ratios are not physically reachable
Several Table 3 rows report VJP cheaper than a standalone forward:
The VJP column times
jax.grad, which dispatches oneapplyRPC followed by onevector_jacobian_productRPC, and most VJP endpoints re-run the forward internally. A VJP measurement is therefore about two forward solves plus one adjoint, which cannot come out below 1× a forward. JAX-FEM and Firedrake sit where this predicts (~2.0–2.4). The sub-1.0 rows should not be reachable. (Credit: @jorgensd, #182 §3.)The current suite confirms this directly. Structural cost at nx=128 (paper resolution) on our CI shows every solver's VJP/forward ratio at 1.7 or above, including the ones the paper puts below 1.0:
Firedrake and JAX-FEM match the paper. FEniCS and TopOpt.jl flip from sub-1.0 to above 1.0 under re-measurement.
3. torch-fem row: re-measured at ~1.6×, not 0.33
We reproduced the torch-fem thermal endpoints in a container (torch-fem 0.6.2, CPU, which is the tesseract's actual numeric path since it hardcodes
_SPARSE_SOLVE_DEVICE="cpu"), instrumenting everysparse_solve:4. TopOpt.jl (0.19) — not yet checked
The listed strategy is "Analytic." For self-adjoint compliance minimisation the λ=−u shortcut means no separate adjoint solve, so a genuinely cheap VJP is plausible. It still needs a forward solution, so the same "cannot be far below a forward" question applies. Not yet independently measured.
Plan / tracking
tesseract_core'sjax_recipesresidual cache could make it "marginal cost of a gradient" if enabled suite-wide (see fix: tune FEniCS Tesseracts (better linear solver, avoid redundant forward calls, only setup once) #182 §3)References