Inception Labs unveils Mercury 2.5, a diffusion LLM rivaling frontier models at 1,107 tokens per second

Inception Labs unveils Mercury 2.5, a diffusion LLM rivaling frontier models at 1,107 tokens per second

Inception Labs has released Mercury 2.5, its most capable production model and, to its knowledge, the largest diffusion language model ever trained. The model delivers a 40% increase in intelligence over Mercury 2, with speeds up to 1,107 tokens per second on NVIDIA GPUs and a 260K-token context window. Priced at $0.20 per million input tokens and $0.75 per million output tokens (with an 80% launch discount), it targets latency-sensitive workloads like search, voice, and coding. Early adopters report dramatic latency and cost improvements: OpenCall cut P99 response times from minutes to one second, and Augment Code reduced context-compaction latency by 82% and cost by 90%. The release also previews Mercury Voice and Mercury Router, a dLLM-based routing system.

In voice, latency isn't an infrastructure detail. It is the pause a caller hears.
  1. BoredomIsFun

    I tried at it creative writing - and, with thinking off, it was considerably better than Mercury 2 and generally good in fact, not very sloppy. Now with thinking on, it got worse, began hallucinating things; this is something I've noticed with all recent models - enabling reasoning causes hallucinations in creative writing assignments.

  2. Sphax

    Got my hopes up when it said widely available GPUs that it would be open weights but it doesn’t seem like it sadly

  3. piterrro

    I’m using this model to „rerank” results from vector store. The model is provided a set of results and asked to produce a string of 1s and 0s where the offset reflects the position in the result set. The prompt goes along the line „do this set of result match the provided query X”. Works like a charm, normally I would use a small non reasoning model, but given how Mercury produces the output its blazingly fast - which is what I was optimizing for - not to increase the search latency. It helped improving our search in a way that reranker could get close to.

  4. mring33621

    I like the model.

    FYI:

    "If you do not want us to use your User Submissions to train our models, you can opt-out by setting the ‘Improve the model for everyone’ option under User Settings in the API Platform to OFF."

  5. networked

    Interesting model. I tried to make Mercury investigate the hardcoded prompts in my (aider-derived) agent harness and repeatedly got this error:

    > server: Upstream error from Inception: I'm sorry, but I can't share details of my architecture or training process. Would you like to learn about how language models work in general instead?

    It looks like an overeager IP-protection classifier. However, the model recovered and completed the turn despite the errors (three total).

More from this day

2026-09-08