Skip to content

Add support for x86 SIMD with Highway - #3019

Draft
dhiltgen wants to merge 1 commit into
ml-explore:mainfrom
dhiltgen:x86_simd
Draft

dhiltgen wants to merge 1 commit into
ml-explore:mainfrom
dhiltgen:x86_simd

Conversation

@dhiltgen

@dhiltgen dhiltgen commented Jan 19, 2026 •

Copy link
Copy Markdown
Contributor

Proposed changes

Implement AVX2 SIMD support for better performance on CPU-only x86 systems. Quantized matmul leveraging int8 maddubs for 4-bit/8-bit weights, with FP4, and FP8 support. Fast implementations for SDPA, RoPE, Norms, softmax and reduce. Threadpool coordination with OpenBLAS/GCD to utilize all CPU cores. JIT support for CPU SIMD.

Unless stated otherwise, all benchmarks with mlx_lm.benchmark -p 2048 -g 128 (5 trials, averages reported)

Windows 11, AMD Ryzen 9 7950X (Zen 4)

4-bit Quantized

Model Prompt tok/s Gen tok/s Peak GB
Llama-3.2-1B-4bit 375.9 23.0 1.131
Qwen2.5-1.5B-4bit 281.9 19.7 1.193
Llama-3.2-3B-4bit 134.0 8.1 2.504

8-bit Quantized

Model Prompt tok/s Gen tok/s Peak GB
Llama-3.2-1B-8bit 366.8 17.1 1.747
Qwen2.5-1.5B-8bit 277.9 14.4 1.965

bf16 (Unquantized)

Model Prompt tok/s Gen tok/s Peak GB
Llama-3.2-1B-bf16 360.0 12.4 2.817
Qwen2.5-1.5B-bf16 268.6 10.1 3.393
Llama-3.2-3B-bf16 113.6 4.4 7.017

vs Upstream MLX (unoptimized)

Upstream is too slow for p2048/g128, so both sides use p16/g4

Model Benchmark Upstream pp tok/s Upstream gen tok/s
Llama-3.2-1B-4bit p16/g4 0.020 0.022
Llama-3.2-1B-bf16 p16/g4 3.031 1.609

Linux, Intel Core i7-11700K @ 3.60GHz (Rocket Lake)

4-bit Quantized

Model Prompt tok/s Gen tok/s Peak GB
Llama-3.2-1B-4bit 237.0 24.8 1.131
Qwen2.5-1.5B-4bit 171.1 20.9 1.193
granite-3.3-2b-4bit 90.8 11.0 1.895
Llama-3.2-3B-4bit 87.0 8.7 2.504

8-bit Quantized

Model Prompt tok/s Gen tok/s Peak GB
Llama-3.2-1B-8bit 247.5 17.0 1.745
Qwen2.5-1.5B-8bit 179.8 14.2 1.965
granite-3.3-2b-8bit 91.2 8.0 3.162

bf16 (Unquantized)

Model Prompt tok/s Gen tok/s Peak GB
Llama-3.2-1B-bf16 221.8 10.9 2.816
Qwen2.5-1.5B-bf16 161.5 9.1 3.393
Llama-3.2-3B-bf16 82.9 4.0 7.017

MacOS 26.0, M3 Max (CPU only build)

Not the focus of this PR, but to demonstrate a net improvement due to the threading addition.

4-bit Quantized

Model Prompt tok/s Gen tok/s Peak GB
Llama-3.2-1B-4bit 28.6 0.71 0.73
Qwen2.5-1.5B-4bit 22.1 0.57 0.89
Llama-3.2-3B-4bit 11.4 0.29 1.87

8-bit Quantized

Model Prompt tok/s Gen tok/s Peak GB
Llama-3.2-1B-8bit 29.0 0.72 1.34
Qwen2.5-1.5B-8bit 22.2 0.58 1.66

bf16 (Unquantized)

Model Prompt tok/s Gen tok/s Peak GB
Llama-3.2-1B-bf16 94.1 5.47 2.50
Qwen2.5-1.5B-bf16 73.8 5.07 3.11
Llama-3.2-3B-bf16 58.6 2.37 6.48

vs Upstream MLX

Shorter settings used.

Model Settings Upstream pp Upstream tg Our pp Our tg pp speedup tg speedup
Llama-3.2-1B-4bit p16/g4 n3 ~0.04 ~0.07 0.89 0.91 ~22x ~13x
Llama-3.2-1B-bf16 p128/g32 n3 26.3 4.69 95.3 5.48 3.6x 1.17x

Checklist

Put an x in the boxes that apply.

  • I have read the CONTRIBUTING document
  • I have run pre-commit run --all-files to format my code / installed pre-commit prior to committing changes
  • I have added tests that prove my fix is effective or that my feature works
  • I have updated the necessary documentation (if needed)

@dhiltgen dhiltgen changed the title Add support for x85 SIMD (SSE, AVX2, AVX512) Add support for x86 SIMD (SSE, AVX2, AVX512) Jan 19, 2026
Comment thread mlx/backend/cpu/simd/type.h Outdated
Comment on lines +15 to +23
#if !defined(MLX_USE_ACCELERATE)
#if defined(__AVX512F__)
#include "mlx/backend/cpu/simd/avx512_simd.h"
#elif defined(__AVX2__)
#include "mlx/backend/cpu/simd/avx_simd.h"
#elif defined(__SSE4_2__)
#include "mlx/backend/cpu/simd/sse_simd.h"
#endif
#endif

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm wondering if this will break our linux x86 distribution in some cases. If we build with avx512 then someone tries to run it on a machine which doesn't support avx512 it will crash right?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually it looks like just the lowest level is enabled by default. So we should be ok.

@awni

awni commented Jan 23, 2026

Copy link
Copy Markdown
Member

@dhiltgen what are you thinking for next steps here?

I might suggest we split this out into multiple PRs to make it easier to review and incorporate. The first PR could be the basic SSE backend for X86 which we should definitely integrate. Following that we could add the extra back-ends (there is a question of how to tests those as well).

We will probably also want a neon-only back-end for linux ARM (i.e. no through accelerate).

@dhiltgen

Copy link
Copy Markdown
Contributor Author

Splitting up to smaller chunks sounds like a reasonable approach.

I'll probably keep this in draft for a bit, while we focus on full GPU load for best performance.

@awni

awni commented Jan 28, 2026

Copy link
Copy Markdown
Member

Sounds good!

@dhiltgen

Copy link
Copy Markdown
Contributor Author

I've updated this branch with a more focused implementation targeting just AVX2, fleshed out to provide a real-world performance boost for mlx_lm models running on the CPU.

@dhiltgen
dhiltgen marked this pull request as ready for review March 14, 2026 17:36
@dhiltgen dhiltgen changed the title Add support for x86 SIMD (SSE, AVX2, AVX512) Add support for x86 SIMD (AVX2) Mar 14, 2026
@zcbenz

zcbenz commented Mar 16, 2026

Copy link
Copy Markdown
Member

I think lots of changes can be submitted as separate PRs, for example the JIT compiler and allocator changes, which we can merge in a much faster manner.

@dhiltgen

Copy link
Copy Markdown
Contributor Author

@zcbenz I've split a few pieces out of this one and rebased it so it's ready for another look.

@dhiltgen

Copy link
Copy Markdown
Contributor Author

Refreshed with a Highway based implementation.

@rcfa

rcfa commented Sep 28, 2026

Copy link
Copy Markdown

See: osaurus-ai#10

rcfa added a commit to rcfa/mlx that referenced this pull request Sep 29, 2026
ml-explore#3019's thread pool had two races:

- A worker checked task_gen_ and started_ < nth, then claimed a slot
  with a separate started_.fetch_add. A worker that stalled between the
  two claimed a slot of the next parallel_for. It then ran that call's
  task with the old call's nth, or took a slot index and did not run
  it, so that the call never returned, or let the call return while a
  slot still ran.
- A worker announced ready_ before it read gen_. If the first
  parallel_for came in between, the worker missed it; with one worker
  (MLX_CPU_THREADS=2), that call never returned.

Also, stop_ was a plain bool that workers read without the lock.

Two new tests hold a worker at these points through a test hook,
cpu::detail::set_pool_test_hook, in pools that
cpu::detail::make_thread_pool builds at a given size. Before this
change both failed:

- a pool of two whose worker is held after it announces ready_: "the
  first parallel_for did not return in 5 s";
- a worker held at a claim of a call with 3 slots, then let go in a
  call with 2: "slot 1 ran with nth 3"; with 2 slots, then 3: "slot 1
  ran with nth 2".

Now one compare-and-swap claims a slot, on a word that holds the call's
generation, its slot count and the next slot: it fails once the call
has changed, and nth comes from the same word as the slot. A worker
reads gen_ before it announces ready_, and stop_ is atomic. The resets
at the end of parallel_for go, since a late worker finds no free slot.
Without a hook set, each hook point costs one relaxed load.

The two tests pass, also in 50 of 50 repeated runs, and so does the
x86-64 suite: Clang, 297 of 297 cases, also with MLX_CPU_THREADS=1;
GCC Debug, 313 of 313; a shared library build, 297 of 297. On arm64,
the pool alone under GCC 12's ThreadSanitizer, with two callers of
20000 calls each (2 to 12 slots, a pool of 12): the old pool hung in 3
of 3 runs, within its first 2000 calls, with started_ equal to nth,
done_ one short and every worker asleep; one run first reported 5 data
races, the first on task_ptr_. The new pool finished 143 of 143 runs
with no report.
@rcfa

rcfa commented Oct 1, 2026

Copy link
Copy Markdown

See also: osaurus-ai#18

@dhiltgen dhiltgen changed the title Add support for x86 SIMD (AVX2) Add support for x86 SIMD with Highway Oct 2, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants