Repository navigation
perf(speculative): re-measure the batch-capable B=1 MTP hardware gate on M1 Ultra and M3 Ultra and adjust mtp_b1_default #1217
Description
Activity
- addedtype:performancePerformance improvementsPerformance improvementspriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked onarea:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)area:benchmarkBenchmark harness and performance measurement (bench_*.sh, /update-benchmarks)Benchmark harness and performance measurement (bench_*.sh, /update-benchmarks)platform:macosmacOS (Apple Silicon) specificmacOS (Apple Silicon) specific
on Aug 19, 2026 PR #1254 covers the M3 Ultra half of this issue and adjusts the gate. It intentionally carries no closing keyword, so this issue stays open for the remaining hardware-gated work.
Done
- Gemma 4 31B + bf16 assistant, B=1, M3 Ultra, 2026-08-20, current main, under the feat(bench): measure speculative decoding properly, and record what it measures #1215 protocol: 2.65x (enumeration), 2.41x (source code), 1.95x (prose), every row inside a 1.0% spread. Round cost 1.51 to 1.52 classic steps across all three, against 2.96 to 3.99 tokens emitted per verify. Width sweep on the code row fits
0.83 + 0.170 K, largest residual 0.06, peak at width 5 with width 4 tied inside its spread. - Gate adjusted per the first branch of the decision matrix:
mtp_b1_defaultnow readsAppleSiliconGen::wide_quantized_projections(theuse_qmv_widesplit at Apple GPU generation 15) instead ofhas_neural_accelerator. Docstring, unit tests,docs/benchmarks.mdanddocs/environment-variables.mdupdated. - Exercised through the real dispatch path (
Scheduler::mtp_b1_should_run) on real checkpoints on M3 Ultra:maindeclines, the branch runs the burst (block 4, 80 tokens over 21 rounds, acceptance 0.921),MLXCEL_ENABLE_MTP_B1=0still declines. - Both bench scripts gained a
gemma31bcase. Neither could reach this pairing before, which is most of why the founding numbers went stale unnoticed. - Record:
docs/benchmark_results/mtp-b1-gate-m3ultra-2026-08-20.md.
Two things worth knowing before the remaining work
The M3 Ultra Qwen rows this issue asks for already exist. They were measured on current main under this protocol in #1215 and are in
docs/benchmarks.md(1.67x at block 3, plus the full width sweep). They were not re-run.The offline
generatepath never consults this gate.mtp_b1_defaulthas exactly one caller,Scheduler::mtp_b1_should_run, andMtpPolicyis built only in the server worker. So theMLXCEL_ENABLE_MTP_B1=1andMLXCEL_MTP_ADAPTIVE=0this issue prescribes are inert in the bench harness; the harness runs the burst unconditionally, which is what makes it the right instrument, but any acceptance check on the gate itself has to go through the server.Remaining, and a prediction to test
M1 Ultra rows for this pairing, and an M5 Max re-measurement. No generation-13 host was available here.
Generation 13 was left declining, and the width sweep is why that is a conclusion rather than an omission. The tempting inference is to take M1 Ultra's published block-4 round cost of 2.71, note this pairing emits 2.96 to 3.99 tokens per verify, and conclude it now clears break-even. That transfers a round cost between pairings with different drafter dtypes, and the sweep shows their slopes differ by 1.9x (
0.83 + 0.170 Khere against the 12B pairing's1.14 + 0.090 Kon the same host; the lines merely cross at K = 4).Carrying the slope ratio onto generation 13's
1.35 + 0.346 Kinstead predicts a block-4 round near 3.6 classic steps on M1 Ultra, so roughly 0.82x to 1.11x across these three prompts. That extrapolates across a pairing and a generation at once, so it is stated as falsifiable rather than believed. An M1 Ultra run withmodels/gemma-4-31b-it-4bitandmodels/gemma-4-31b-it-assistant-bf16plusscripts/bench_speculative.sh gemma31bsettles it in one sweep. If it comes out well above break-even, the generation-13 half of the gate should move too.Also worth re-running when a host is free: M5 Max's ~1.2 to 1.4x is the founding, pre-#1203 figure, and the M3 Ultra rows above now exceed it.
- Gemma 4 31B + bf16 assistant, B=1, M3 Ultra, 2026-08-20, current main, under the feat(bench): measure speculative decoding properly, and record what it measures #1215 protocol: 2.65x (enumeration), 2.41x (source code), 1.95x (prose), every row inside a 1.0% spread. Round cost 1.51 to 1.52 classic steps across all three, against 2.96 to 3.99 tokens emitted per verify. Width sweep on the code row fits
- addedstatus:in-progressCurrently being worked onCurrently being worked onand removedstatus:readyReady to be worked onReady to be worked on
on Aug 19, 2026 - addedstatus:readyReady to be worked onReady to be worked onand removedstatus:in-progressCurrently being worked onCurrently being worked on
on Sep 1, 2026
Summary
The static per-hardware gate for the B=1 MTP burst,
mtp_b1_defaultinsrc/server/batch/speculative_burst.rs, decides as follows: anMLXCEL_ENABLE_MTP_B1override wins in both directions; otherwise non-batchable targets always run the burst, and batch-capable targets run it only whenhas_neural_acceleratoris true (an M5+ chip-generation proxy frommlxcel_core::hardware). The pre-M5 decline for batch-capable targets rests on measurements taken before several substantial MTP changes landed, and M3 Ultra has never been measured on the batch-capable pairing at all: the binary NA proxy lumps it with M1 Ultra even though every recent measurement places it much closer to M5 Max. This issue asks to re-measure the batch-capable B=1 pairing on M1 Ultra and M3 Ultra on current main and to adjust the gate according to the results.Current behavior
mtp_b1_default(src/server/batch/speculative_burst.rs, pure decision core ofmtp_b1_burst_enabled): batch-capable target + no neural accelerator = decline to classic decode. Its docstring records the founding measurements: batch-capable 31B + bf16 assistant at ~1.2 to 1.4x on M5 Max, but a consistent 0.75 to 0.96x regression on M1 Ultra (four greedy 160-token prompts), which is why pre-M5 defaults to classic. That policy dates to perf: validate B=1 MTP default-on on lower-bandwidth Apple Silicon (M1/M2) #165.Scheduler::mtp_b1_should_run(src/server/batch/scheduler.rs) when no adaptive policy is attached, i.e. wheneverMLXCEL_MTP_ADAPTIVE=0.MtpPolicy::static_default(src/server/batch/mtp_policy.rs) calls the samemtp_b1_defaultto resolve an ambiguous profiling window. So a stale static default still steers real decisions even in the adaptive path.Why the founding evidence is stale
qmv_wide), perf(speculative): quantize the MTP drafter's projections at load #1203 (drafter projections quantized at load), fix(speculative): report the block width the MTP round loop actually used #1208 (report the effective block width the round loop actually used), feat(bench): measure speculative decoding properly, and record what it measures #1215 (benchmark methodology: GPU contention guard, round-cost protocol,scripts/with_indexers_paused.sh). In particular perf(speculative): quantize the MTP drafter's projections at load #1203 is documented indocs/benchmarks.mdas the change that moved the Qwen pairing on M1 Ultra from 0.59 to 0.70x up to break-even (docs/benchmark_results/qwen38-mtp-m1ultra-2026-08-16.md), so its effect on the Gemma 4 31B + bf16 assistant pairing on pre-M5 hosts is plausibly material and has not been measured.docs/benchmarks.md(measured on the non-batchable Gemma 4 12B pairing) puts M3 Ultra at a break-even of ~1.51 emitted tokens per verify at block 4, far closer to M5 Max's ~1.28 than to M1 Ultra's ~2.71, and M3 Ultra measures 1.74 to 2.61x on the B=1 12B pairing.use_qmv_widesplit (documented insrc/models/speculative_exactness.rs): generation 15+ (M3, M4, M5) runs a quantized projection at M >= 2 as one wide pass, generation 13 (M1, M2) runs the block as narrow per-position passes.has_neural_accelerator(M5+ only) is a coarser proxy than the mechanism, and it misclassifies M3 Ultra to the slow side.What to measure
On both M1 Ultra and M3 Ultra, on current main, following the #1215 protocol (
scripts/bench_speculative.shunderscripts/with_indexers_paused.sh, spread and contention guards respected, adaptive profiling window discarded, host and prompt recorded on every row):MLXCEL_ENABLE_MTP_B1=1andMLXCEL_MTP_ADAPTIVE=0so the static gate and the policy do not interfere with the measurement. Run the three standard prompts (enumeration, source code, prose), plus a block-width sweep (scripts/bench_block_width.sh) at least on the code row.docs/benchmarks.md; the target is batch-capable so the same gate governs it).Checkpoints needed:
gemma-4-31b-it-4bit,gemma-4-31B-it-assistant-bf16(already wired intospeculative_benchREACHABLE_PAIRINGS), and the Qwen pairing checkpoints on the M3 Ultra host.Adjusting the gate (decision matrix)
has_neural_acceleratorpredicate inmtp_b1_defaultwith the finer discriminator the measurements support, e.g. Apple GPU generation >= 15 (theuse_qmv_widecapability the exactness probe already reads), keeping generation 13 declined.docs/benchmarks.mdwith the new dated rows so the next reader knows the evidence is current.mtp_b1_defaultunit tests (it is the pure test seam for this decision), update the docstring measurements, and keep the note that the static default also resolves adaptive-policy ambiguity, so its correctness matters even with the adaptive path on.Acceptance criteria
docs/benchmark_results/and reflected indocs/benchmarks.md, following the feat(bench): measure speculative decoding properly, and record what it measures #1215 protocol (guards, host and prompt recorded, profiling window discarded).mtp_b1_defaultadjusted per the decision matrix (or explicitly confirmed unchanged with the dated evidence), with its docstring updated to the new measurements.mtp_b1_defaultupdated to the new predicate and passing.Scheduler::mtp_b1_should_run) on at least one real checkpoint per affected host, not only through unit tests.Notes
MLXCEL_ENABLE_MTP_B1remains the both-directions override and is unaffected by this issue.