Skip to content

Running list: potential discrepancies in preprint timings #183

Description

@dionhaefner

Summary

Follow-up work on the differentiable-solver benchmarks (dolfinx-adjoint#81, #179, #182) has surfaced several timing discrepancies between Table 3 of the paper (arXiv:2606.27895) and what the current benchmark suite reproduces. This issue collects them in one place and tracks their resolution.

These look like artifacts of the earlier benchmark harness and the environment the paper's numbers were produced in, not fundamental problems with the solvers or the physics. The forward and gradient accuracy results are unaffected. Only the wall-clock cost columns are in question. We intend to re-run the affected timing measurements on the current suite and hardware, incorporating the first wave of third-party solver fixes, and update the table accordingly.

The paper is still at the preprint stage, so correcting these now is straightforward. To be clear, this is a one-time reconciliation while the work is pre-publication, not an open-ended commitment to chase every measurement indefinitely.

Discrepancies

1. FEniCS structural forward time (~140 s) does not reproduce

Table 3 lists the legacy-FEniCS structural forward at ~140 s (N=128). On the current suite (#179 extended the cost sweep to the paper's 128×2×64 resolution), FEniCS forward is ~5.5 s and VJP ~16.3 s on CI. After #182 (UMFPACK→mumps plus tape caching) it drops to ~3.7 s / ~11.1 s. That is a ~25–38× gap from the published number.

  • The FEniCS solver code is essentially unchanged since it was written, so this is not a solver regression that was later fixed.
  • Instrumentation (see fix: tune FEniCS Tesseracts (better linear solver, avoid redundant forward calls, only setup once) #182) shows ~99.6% of the structural time is the linear solve, with the adjoint costing ~1.0× the forward. There is no path from ~11 s to ~140 s on this solver.
  • The most likely cause is the paper-repro environment (thread count, solver defaults, no warm-tape reuse) rather than the solver. An end-to-end re-run is needed to confirm.

2. Sub-1.0 VJP/forward ratios are not physically reachable

Several Table 3 rows report VJP cheaper than a standalone forward:

row Fwd VJP ratio
FEniCS (H) 60 s 47 s 0.78
FEniCS (S) 140 s 86 s 0.61
torch-fem (H) 33 s 11 s 0.33
TopOpt.jl (S) 78 s 15 s 0.19

The VJP column times jax.grad, which dispatches one apply RPC followed by one vector_jacobian_product RPC, and most VJP endpoints re-run the forward internally. A VJP measurement is therefore about two forward solves plus one adjoint, which cannot come out below 1× a forward. JAX-FEM and Firedrake sit where this predicts (~2.0–2.4). The sub-1.0 rows should not be reachable. (Credit: @jorgensd, #182 §3.)

The current suite confirms this directly. Structural cost at nx=128 (paper resolution) on our CI shows every solver's VJP/forward ratio at 1.7 or above, including the ones the paper puts below 1.0:

solver (structural, nx=128) fwd vjp vjp/fwd Table 3 ratio
FEniCS 5.47 s 16.31 s 2.98 0.61
Firedrake 5.07 s 12.64 s 2.49 2.4
JAX-FEM 13.51 s 38.48 s 2.85 2.3
TopOpt.jl 0.63 s 1.06 s 1.69 0.19

Firedrake and JAX-FEM match the paper. FEniCS and TopOpt.jl flip from sub-1.0 to above 1.0 under re-measurement.

3. torch-fem row: re-measured at ~1.6×, not 0.33

We reproduced the torch-fem thermal endpoints in a container (torch-fem 0.6.2, CPU, which is the tesseract's actual numeric path since it hardcodes _SPARSE_SOLVE_DEVICE="cpu"), instrumenting every sparse_solve:

N DOFs apply VJP ratio solves/apply solves/VJP
64 4,290 ~0.4 s ~0.6–1.1 s 1.4–2.5 1 2
128 16,770 ~2.4–3.5 s ~3.4–4.8 s ~1.4 2 3
256 66,306 ~7–9 s ~12–14.5 s ~1.60 2 3
  • VJP/forward is ~1.6 at the paper's N=256, not 0.33. The VJP always does one extra adjoint solve on top of the forward, exactly as the source implies. This is consistent with @jorgensd's in-process ~1.5.
  • The preconditioner reuse the paper describes is real. The adjoint reuses the AMG hierarchy, making it ~0.77× a forward. That reduces the adjoint cost but cannot invert the ratio below 1.
  • Our absolute forward at N=256 is single-digit seconds against the published 33 s. Like the FEniCS 140 s, that large gap points to the earlier environment, so the torch-fem row is worth re-running end-to-end rather than just checking its ratio.
  • Caveat on our numbers: CPU, in-process (not over the tesseract RPC), on a laptop, so absolute times carry run-to-run variance. The ratio stays above 1 at every size across repeats.

4. TopOpt.jl (0.19) — not yet checked

The listed strategy is "Analytic." For self-adjoint compliance minimisation the λ=−u shortcut means no separate adjoint solve, so a genuinely cheap VJP is plausible. It still needs a forward solution, so the same "cannot be far below a forward" question applies. Not yet independently measured.

Plan / tracking

References

Activity

  1. changed the title [-]Reconcile Table 3 timing discrepancies with the current benchmark suite[/-] [+]Running list: potential discrepancies in preprint timings[/+] on Sep 18, 2026
  2. added theissue type on Oct 1, 2026
  3. meyer-nils commented on Oct 6, 2026

    @meyer-nils
    Contributor

    @dionhaefner
    While you are at it, I took a quick look at the torch-fem related issues here and noted the following things:

    • feat(solver): update torch-fem to 0.12.1, solve with CG+Jacobi on GPU #187 bumped the torch-fem version to 0.12.1. If you run the benchmarks again, it should ideally use that more recent version.
    • The statement direct (spsolve) in Table 2 is not strictly correct. On GPU, the thermal example uses CG + Jacobi (more iterations are faster than spending time on an advanced preconditioner here). On CPU, the thermal example dispatches based on DOF count: Below 10_000, it's direct (spsolve), above it's CG + AMG. That's why there was a jump in the table and the paper's N=256 was never a direct solve. I suggest the GPU CG-Jacobi path with a clear indication of GPU usage.
    • feat(solver): Add torch-fem backend for structural-mesh #188 added torch-fem as an additional structural backend, so it could extend the torch-fem domains to H and S in the paper.
  4. dionhaefner commented on Oct 6, 2026

    @dionhaefner
    ContributorAuthor

    Thanks @meyer-nils ! We'll make sure to re-run the whole suite end-to-end before this gets published anywhere.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions