Skip to content

Add simd_swizzle! macro for convenient compile-time swizzles - #391

Open
Shnatsel wants to merge 3 commits into
linebender:mainfrom
Shnatsel:const-swizzle
Open

Shnatsel wants to merge 3 commits into
linebender:mainfrom
Shnatsel:const-swizzle

Conversation

@Shnatsel

@Shnatsel Shnatsel commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Mirrors std::simd API on the surface but not in the implementation.

Does not (yet?) support resizing vectors, unlike std::simd; the input and output are always the same width, just swizzled.

#354 will require a separate concat_swizzle! macro, whereas std::simd subsumes this into the resizing functionality.

In exchange we gain the ability to use simd_swizzle! in generic contexts. Native-width vectors can now use simd_swizzle! and get a subslice of a predefined index, or even call a const fn that computes the indices without hardcoding them at all.

All the internals are #[doc(hidden)] and unstable. Only the macro is the public API so we could evolve it without breaking changes.

An MVP would be just the function to map indices, but the macro is much more convenient and doesn't add much complexity.

@simonask

simonask commented Sep 25, 2026 •

Copy link
Copy Markdown

Currently traveling, but here’s my feedback (I believe I motivated this from a Reddit comment):

  • It should be possible to do this with a const fn rather than a macro. In my code, I have a const fn that translates [u8; LANES] to [u8; BYTES] (so lanes [N, …] becomes bytes [N*W+0, N*W+1, N*W+2, …], where W is the lane width in bytes).
  • Naming: You could consider naming this operation “shuffle” instead of “swizzle” to disambiguate in the API, and hint at the resulting intrinsic. .NET and other libraries use Shuffle() for the “precise” variant, and ShuffleNative() for the implementation-defined out-of-bounds behavior.
  • In my code I have a Shuffle<const N: usize> trait for convenience, where N is the number of lanes in the result, which I have also implemented for double the number of lanes of each vector type. So f32x4 implements both Shuffle<4> and Shuffle<8>, where the latter produces an f32x8.

The Shuffle trait looks like this:

pub trait Shuffle<const N: usize> {
    type Output;
    fn shuffle(self, lanes: [u8; N]) -> Self::Output;
}

Now, I wouldn’t know how to express that for variable-width vector types.

I can share my code next week when I’m no longer traveling.

@Shnatsel

Shnatsel commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor Author

It should be possible to do this with a const fn rather than a macro.

Yes, it is a const fn internally, the macro wraps it for convenience. There are three reasons for using the macro:

  1. Calling the const fn directly is a lot of faff, you need to pass the bit width of the type somehow which is inconvenient and error-prone.
  2. The macro approach also matches the std::simd API.
  3. The macro allows us to implement future optimizations; for example, lowering into swizzle_dyn_within_blocks directly and relying less on LLVM to optimize the shuffle, especially for wider-than-native vectors.

Naming: You could consider naming this operation “shuffle” instead of “swizzle” to disambiguate in the API,

The naming follows std::simd. That is what the ecosystem has converged on, with wide also following std::simd naming closely, so we are not going to deviate from the ecosystem standard.

In my code I have a Shuffle<const N: usize> trait for convenience

This is similar to how std::simd implements it, but that makes it impossible to use on hardware-sized vectors the size of which isn't known at compile time. I am taking a different approach here specifically to let it work on hardware-sized types such as f32s.

This was referenced Sep 29, 2026
pthariensflame pushed a commit to pthariensflame/fearless_simd that referenced this pull request Oct 6, 2026
We have enough changes for a release and nothing but
linebender#391 outstanding, so
this seems like a good time to cut one.

---------

Co-authored-by: Laurenz Stampfl <laurenz.stampfl+github@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants