-

Inception Launches Mercury 2.5, the Next Tier of Intelligence for Diffusion LLMs

  • As AI infrastructure costs climb, diffusion LLMs generate tokens in parallel instead of one-at-a-time, making them the fastest and most cost-efficient LLM solution on the market.
  • Diffusion has moved from a research bet to production standard: Since Mercury 2’s launch, enterprise Mercury usage has grown to thousands of developers and dozens of enterprises, powering latency-sensitive workflows across search, voice, and coding agents at scale.
  • Founded by the Stanford, UCLA, and Cornell researchers behind the first diffusion LLM, Inception now ships Mercury 2.5 with more intelligence, lower cost, and the same speedup over traditional models.

REDWOOD CITY, Calif.--(BUSINESS WIRE)--Inception, the company behind the first commercial diffusion large language models (dLLMs), today announced the launch of Mercury 2.5, the most capable dLLM and the fastest reasoning LLM in production. Mercury 2.5 delivers the next tier of intelligence for dLLMs and runs over 1,100 tokens per second in production.

Most LLMs in production today still generate text autoregressively: one token at a time. That approach ties cost and latency directly to reasoning depth, so the more a model has to think, the slower and more expensive it becomes. dLLMs work differently, starting with a rough draft of the output and refining tokens in parallel.

dLLMs have moved from a research bet to a production workhorse. Enterprise Mercury usage is growing by over an order of magnitude, powering search, voice, and coding agents at scale. Inception proved that diffusion was enterprise-ready with Mercury 2, matching Claude Haiku and GPT Mini intelligence at roughly 10x the throughput, and Mercury models have since become the dLLM of choice for teams that can't afford to trade intelligence for speed, or speed for cost. Mercury 2.5 builds directly on that foundation: more intelligence, lower cost, and the same speedup over traditional models.

“Nobody in this industry thinks LLMs can get smarter, faster, and cheaper at the same time. It’s a limitation of autoregressive modeling,” said Stefano Ermon, CEO and co-founder of Inception. “Mercury 2.5 proves the switch to diffusion opens that door.”

Token Efficiency

As AI infrastructure costs climb and data centers strain to keep up with demand, token efficiency is under more scrutiny than it was even a year ago. Every token costs compute, power, and data center capacity, and autoregressive models spend all three generating one token at a time. dLLMs make fundamentally better use of each GPU cycle, generating many tokens per step instead of one, which means more intelligence per dollar. That efficiency runs on widely-available NVIDIA GPUs, not specialized silicon.

Smarter, Not Slower

Mercury 2.5 jumps 10 points in intelligence over Mercury 2, making it more intelligent than any previous dLLM, comparable to cost-optimized frontier models like GPT-5.6 Luna, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Throughput is up to over 1,100 tokens per second, and pricing drops to $0.20 / $0.75 per 1M tokens. At launch, Mercury 2.5 is 80% off at $0.04 / $0.15 per 1M tokens. Mercury 2.5 extends the context window to 260K tokens, up from 128K, and it adds tunable reasoning, native tool use, JSON mode, and parallel tool calls.

“After switching to Mercury, our P99 response time dropped from several minutes to just one second, and P50 dropped from 0.4 seconds to less than 0.2 seconds, faster than any other provider we've seen on the market, including reasoning,” said Oliver Silverstein, Co-founder and CEO of OpenCall.

Where Diffusion Thrives

Mercury dLLMs are already running in production across the workloads that demand responsive inference:

  • Search Agents and RAG Pipelines: One search request can trigger dozens of model calls: plan the search, rewrite queries, rerank results, structure facts, summarize sources, and check the answer. Mercury keeps those calls fast enough to stay inside a single user interaction.
  • Voice Agents: Agents respond almost immediately, even with reasoning turned on, fast enough for a phone call with an agent to feel like a real conversation.
  • Coding: Fast enough that developers can prompt, review, and edit code back-to-back without waiting on the model.

Alongside Mercury 2.5, Inception also announced a preview of Mercury Voice and Mercury Router. Mercury Voice delivers time-to-first-token (TTFT) under 170ms and is a dLLM optimized for voice agents with the tightest latency budgets. Mercury Router understands incoming prompts with a dLLM and routes them to the best models (open and closed models) that offer the best mix of quality, speed, and cost. Interested customers can contact sales@inceptionlabs.ai to get preview access to Mercury Voice and Mercury Router.

Mercury 2.5 is enterprise-ready and available now on Inception’s API, OpenRouter, and Baseten. Get started with 100 million free tokens: https://platform.inceptionlabs.ai/

About Inception

Inception develops diffusion-based large language models (dLLMs) designed for efficient, low-latency AI applications. While traditional autoregressive LLMs generate text sequentially, Inception’s diffusion-based models generate outputs in parallel, enabling faster inference and improved reliability for real-time use cases like search, voice, and coding agents. Inception's Mercury models are available via the Inception API, OpenRouter, and Baseten. Based in Redwood City, California, Inception is backed by Menlo Ventures, Mayfield, Innovation Endeavors, M12 (Microsoft’s venture capital fund), Snowflake Ventures, Databricks Ventures, and individual backers including Andrew Ng, Andrej Karpathy, and Eric Schmidt. For more information, visit www.inceptionlabs.ai.

Contacts

Press Contact
Hannah Lulinski
VSC, on behalf of Inception
inception@vsc.co

Inception


Release Versions

Contacts

Press Contact
Hannah Lulinski
VSC, on behalf of Inception
inception@vsc.co

More News From Inception

Inception Launches Mercury 2, the Fastest Reasoning LLM — 5x Faster Than Leading Speed-Optimized LLMs, with Dramatically Lower Inference Cost

PALO ALTO, Calif.--(BUSINESS WIRE)--Inception, the company behind the first commercial diffusion large language models (dLLMs), today announced the launch of Mercury 2, the fastest reasoning LLM and first reasoning dLLM. Mercury 2 delivers 5x faster performance while reducing the latency and cost barriers that have limited real‑world deployment of reasoning systems. Mercury 2 models are available today via the Inception API. Every major LLM in production today, including GPT, Claude, and Gemini...

Inception Raises $50M to Power Diffusion LLMs, Increasing LLM Speed and Efficiency by up to 10X and Unlocking Real-Time, Accessible AI Applications

PALO ALTO, Calif.--(BUSINESS WIRE)--Inception, the company pioneering diffusion large language models (dLLMs), today announced it has raised $50 million in funding. The round was led by Menlo Ventures, with participation from Mayfield, Innovation Endeavors, NVentures (NVIDIA’s venture capital arm), M12 (Microsoft’s venture capital fund), Snowflake Ventures, and Databricks Investment. Today’s LLMs are painfully slow and expensive. They use a technique called autoregression to generate words sequ...

Inventor of Diffusion Technology Underlying Sora and Midjourney Launches Inception to Bring Advanced Reasoning AI Everywhere, from Wearables to Data Centers

PALO ALTO, Calif.--(BUSINESS WIRE)--Inception today introduced the first-ever commercial-scale diffusion-based large language models (dLLMs), a new approach to AI that significantly improves models’ speed, efficiency, and capabilities. Stemming from research at Stanford, Inception’s dLLMs achieve up to 10x faster inference speeds and 10x lower inference costs while unlocking advanced capabilities in reasoning, controllable generation, and multi-modal data analysis. Inception’s technology enable...
Back to Newsroom