AMD's Matrix Cores Break IEEE 754, So Researchers Built Bit-Exact Models

Accurate Models of AMD Matrix Cores

Matrix multipliers on recent GPUs don't follow IEEE 754, and their undocumented quirks—accumulator width, rounding, subnormal handling—make results irreproducible across devices. Researchers characterize AMD's CDNA 1, CDNA 2, and CDNA 3 matrix cores using targeted test vectors, then build MATLAB models that match hardware bit-for-bit across 10 million random inputs. They also quantify application-level accuracy differences between AMD matrix cores and NVIDIA tensor cores.

As a result, reproducibility of small matrix multiplier results across devices is not possible and cannot be achieved by software control.
  1. erichocean

    If you're an author of this, please follow up with the SME2 cores in Apple Silicon, and the AMX cores in Intel server processors.

  2. peter_d_sherman

    >"Features of matrix multipliers differ across vendors and architectures of the same vendor [...] As a result, reproducibility of small matrix multiplier results [differences] across devices is not possible and cannot be achieved by software control. Implementation details of matrix multipliers are not documented, making it difficult to interpret discrepancies in the computed results."

    I'm guessing (but not knowing) that small subtle differences in matrix multiply across different vendor's architectures (and product generations of an individual vendor's architecture) is responsible for a good portion of software crashes when trying to run a local LLM on a different architecture or with a different stack (ROCm vs. CUDA, for example) than the ones it has been explicitly tested on.

    As such, this marks a rather significant problem for the future, which can basically be stated as:

    There needs to be a standard matrix multiply specification (much like IEEE-754 is/was for floating point operations) that all future vendors of AI accelerators (any GPU, CPU, NPU or IC manufacturer whose circuits implement matmul) adhere to, such that the matmul of one vendor is exactly and precisely compatible with the matmul of another.

    Hardware vendors of course, are free to compete in terms of speed, power efficiency, number of matmul engines on a given piece of silicon, parallelization optimizations, etc., but the basic matmul operation should be exactly and precisely compatible across vendors and […]

More from this day

2026-09-16