Etched Sohu vs. NVIDIA: Transformer ASIC vs. GPU (2026)
Etched Sohu vs. Nvidia: Transformer ASIC vs. GPU (2026) – Spheron Blog

Etched AI's Sohu is a transformer-only ASIC that hard-codes attention into silicon, claiming 500,000 tokens/sec on an 8-chip server for Llama 70B—about 62,500 tokens/sec per chip, versus ~700 tokens/sec for an H100 at batch 1. But Sohu sacrifices all programmability: it can't run vision, diffusion, MoE, or SSM models. With $800M raised, $1B in contracts, and first racks shipping summer 2026, the real question is whether the architectural bet pays off for your workload. This analysis compares Sohu against H100, B200, and Groq's LPU, offering a cost-per-token framework and a decision guide.
Sohu's throughput advantage over GPUs comes from architectural specialization of transformer attention patterns built on top of standard HBM3E, not from a SRAM-based design like Groq.
- tonis2
Even if the Sohu Asic chip does attention part super fast, wont the bottlenecks come from, when this data goes to some next stage ?
I'm just thinking whats the chance that we will actually start using ASIC chips for certain parts of AI inference.
- pyrolistical
Can’t wait for single chip asic qwen3.8 27b
It’s small so should be cheap per chip. And it so much smarter than it ought to be for its size.
Problem is asic take forever to cut and are always months behind the latest open weight sota
- rvz
2 hours and no comments makes me wonder if the Etched chip exists or not.
> Sohu figures are per chip, derived from Etched's published 8-chip server claim of 500,000 tok/s on Llama 70B; not independently verified. (claimed by Etched)
So these are not verified benchmarks and they are all claims. Again is this chip real?