Nari Labs' Qwen3-TTS hits sub-50 ms audio response, 10 RPS on one H100
How We Made a Text-to-Speech Model Respond in Sub-50 ms

Nari Labs' implementation of Qwen3-TTS 1.7B CustomVoice achieves sub-50 ms p95 time-to-first-audio (TTFA) at 10 requests per second on a single NVIDIA H100 SXM, while maintaining real-time playback with zero underruns. The team compares five serving engines, tunes them for low latency, and open-sources both the implementation and benchmark. Key optimizations include a unified scheduler for the three-model architecture, dynamic leading-silence trimming, state-cached incremental decoding for the codec, and CUDA graphs. At $4.29/hour for the H100, this translates to roughly $2 per 1M characters, far cheaper than commercial APIs like ElevenLabs V3 ($100/1M) or Cartesia Sonic 3.5 ($49/1M).
By replacing a host-driven sequence with a fixed GPU program, we lower latency and simplify the execution system.
- toebee
time-to-first-audio (TTFA) is critical for realtime voice applications. open source implementations (e.g. vLLM-Omni, SGLang-Omni) are often too slow for production and can have issues with realtime playback if you push for lower latency. we wanted to fix that.
we optimized qwen3-tts, a popular OSS TTS model, to achieve 34 ms p95 TTFA at 10 requests per second on 1 x H100. we open source the implementation and benchmark, as well as a breakdown of how it was done.
- armcat
Having built my own voice assistant (https://github.com/acatovic/ova) and having tried many other services and models, I feel the real win is when this is on-device, and by "on-device" I mean being very inexpensive to run on a phone, and not H100. I've now been using Pocket TTS which is super fast, and also Chatterbox and Fish Audio S2 Pro (on the Mac/PC), I feel we are so close, yet so far. The quality is amazing, but can we take this to the next level and make it run on mobile? What would it take?
- cfferrys
looks good!
- nowittyusername
This is right up my alley as ive been building a local voice agent for a year now. Ive tried many different models and have a custom implementation for omni voice that ive tuned for over many months. Ive never been able to achieve faster then 200ms ttfa for that model at 24 steps, but the reason is .... quality. I find that there is a lot of room for improvement in many tts models out there by a huge margin. But there is also a quality hard wall that you eventually hit that the tradeoff of faster latency but lower quality is not worth it. When making a really well sounding voice agent quality of voice, cadence, expression, etc... matters a lot. It will be interesting to try this implementation and see if its quality outputs match my expectations, if so great job indeed.
- jmesmith
any plans to make this available on cloudflare ai workers (or similar)? Looks super cool, I'd love to try it!