NumPy in the browser gets a 30x speedup with OpenBLAS
Faster NumPy in the Browser

Running NumPy in the browser used to mean slow, unaccelerated matrix math. Now, thanks to Emscripten-forge linking OpenBLAS in WebAssembly, np.matmul at n=1024 is about 30.92x faster for float32 and 14.90x for float64. The upcoming OpenBLAS 0.3.35, with WASM SIMD kernels from QuantStack, pushes performance further, and an optional Relaxed SIMD build adds another boost on supported engines.
The last mile of a long road: faster NumPy in the browser.
- sharktheone
Very cool too see. So happy we actually have SIMD in WASM
- miohtama
Wouldn't AI rewrite of these legacy Fortran packages like BLAS be easier at this point than trying to keep Fortran alive?
- wangxiaoxiang55
The dynamic-linking detail is the part I find most interesting. On PyPI, NumPy wheels vendor a BLAS snapshot, so a BLAS improvement means waiting for the next NumPy release. Here OpenBLAS is its own conda package, so a JupyterLite / notebook.link environment picks up 0.3.35 by adding a channel and nothing else changes. That's a real payoff of treating wasm32 as a conda platform instead of a wheel target.
Question for the authors: Appendix F puts 0.3.35 at roughly 0.4-0.5x of single-thread linux-64 on GEMM. How much of the remaining gap do you attribute to wasm codegen vs. the missing threads? And is a pthreads build on the roadmap at all? A lot of JupyterLite deployments are static hosting (GitHub Pages and the like) where you can't set the COOP/COEP headers SharedArrayBuffer needs, so I'd guess single-thread stays the default for a long time regardless.