Rust's SIMD ecosystem finally grows up

The state of SIMD in Rust in 2026

Sergey "Shnatsel" Davidoff surveys the state of SIMD in Rust for 2026, now as a maintainer of Fearless SIMD. He compares five approaches: std::simd, fearless_simd, wide, pulp, and macerator, across multiversioning, portability, and instruction set support. The takeaway: automatic vectorization remains unreliable, std::simd is nightly-only and incomplete, and fearless_simd emerges as the most complete all-in-one solution.

Floats are weird. Even something as trivial as summing an array of floats with reasonable precision gets surprisingly involved.
  1. ack_complete

    AArch64 definitely has a much more comprehensive baseline than x86-64, but there are some optional extensions that are situationally impactful, including the Crypto extension and some of the newer accumulation / dot product instructions. And unlike Intel, ARM has no portable equivalent to CPUID for querying feature flags and is terrible at documenting which intrinsics require specific FEAT_* flags.

    The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs. For basically any mainstream x86 CPU, it's trivial to find documentation listing what ISA level it supports and general execution widths and latencies for common operations. For the majority of ARM CPUs, there's absolutely nothing. ARM only has optimization guides for selected Cortex cores, and NVIDIA published info for their Olympus core. But execution details had to be reverse engineered for Apple M1, and there is nothing for Oryon. This is especially bad for in-order cores, which unfortunately is still relevant because new CPUs are still being shipped with in-order efficiency cores.

  2. tancop

    > ... you can just assert that they’re all recent enough to at least have AVX2 that was introduced over 10 years ago, and have the program crash or misbehave if it ever runs on anything without AVX2

    > However, if you are distributing the binaries for other people to run, that’s not really an option.

    This all depends on what kind of software you're making. A lot of games set their requirements about 5 generations back, like FC 27 where the minimum is a Ryzen 1600. That lets them use AVX2 unconditionally and prevent complaints from users who tried to run it with a super old CPU.

    Then you get whole Linux distros like CachyOS and Clear (RIP) that rebuild the world for each architecture level and have them as separate variants. I think it still counts as binaries for other people.

  3. sharktheone

    I am hoping for portable SIMD so much. But I still think that often a manually rolled SIMD will be faster.

    Also the state of SIMD in Cranelift is also very WIP. They pretty much just support a subset of 128bit vectors with some rare exceptions.

  4. the__alchemist

    I am using my own lib, `lin_alg`, which apes core_simd for floating point values, and extends the concept to vectors and quaternions. I will eventually replace the floating point portions with core::simd upon its arrival in stable Rust.

    Downside: It's currently x86 only.

  5. Archit3ch

    Hot take: there is no portable SIMD.

    You can either have performance (=write manual ASM for each platform), or portability, but not both.

    What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)

More from this day

2026-09-27