-

Cerebras Helps Power OpenAI’s Open Model at World-Record Inference Speeds: gpt-oss-120B Delivers Frontier Reasoning for All

Industry leading performance for OpenAI gpt-oss-120B at 3,000 Tokens Per Sec

SUNNYVALE, Calif. & SAN FRANCISCO--(BUSINESS WIRE)--Cerebras Systems today announced inference support for gpt-oss-120B, OpenAI’s first open-weight reasoning model, now running at record-breaking inference speeds on the Cerebras AI Inference Cloud. Purpose-built for complex challenges in math, science, and code, this 120B-parameter model achieves intelligence on par with top proprietary models like Gemini 2.5 Flash and Claude Opus 4—while delivering unmatched speed, cost efficiency, and openness.

"Through deployment partners like Cerebras, we're together able to provide powerful, flexible tools that make it easier than ever to build, innovate, and scale," said Dmitry Pimenov, product lead at OpenAI.

Share

For the first time, an OpenAI model leverages Cerebras’ wafer-scale AI infrastructure to run full-model inference. By eliminating GPU memory bandwidth bottlenecks and communication overhead, Cerebras wafer-scale AI inference delivered a world-record 3,000 tokens per second output speed — a major advance in responsiveness for high-intelligence AI.

“OpenAI’s open-weight reasoning model release is a defining moment for the AI community,” said Andrew Feldman, CEO and co-founder of Cerebras. “With gpt-oss-120B, we’re not just breaking speed records—we’re redefining what’s possible. OpenAI on Cerebras delivers frontier intelligence with blistering performance, lower cost, full openness, and plug-and-play ease of use. It’s the ultimate AI platform: smart, fast, affordable, easy to use, and fully open.”

At over 3,000 tokens/second, organizations will be able to use Cerebras-powered gpt-oss-120B to build live coding assistants, instant large document Q&A and summarization, and fast agentic research chains. These high-intelligence AI reasoning use cases have long wait times on proprietary models running on GPUs – that lag is now dramatically reduced with gpt-oss-120B on Cerebras.

Developers can swap their existing OpenAI endpoints for Cerebras in 15 seconds. No refactoring. No migration headaches. Just instant access to the highest performance and quality gpt-oss-120B models running on the Cerebras Cloud.

The open-weight Apache 2.0 license from OpenAI gives users full control to fine-tune for their domain, deploy on-prem for sensitive or regulated data, or move freely across clouds.

"Our open models let developers—from solo builders to large enterprise teams—run and customize AI on their own infrastructure, unlocking new possibilities across industries and use cases," said Dmitry Pimenov, product lead at OpenAI. "Through deployment partners like Cerebras, we're together able to provide powerful, flexible tools that make it easier than ever to build, innovate, and scale."

Experience the fastest AI inference today:
Developers and enterprises can now access gpt-oss-120B on the Cerebras Cloud with a free API key (cerebras.ai/openai).

About Cerebras Systems

Cerebras Systems is a team of pioneering computer architects, computer scientists, deep learning researchers, and engineers of all types. We have come together to accelerate generative AI by building from the ground up a new class of AI supercomputer. Our flagship product, the CS-3 system, is powered by the world’s largest and fastest commercially available AI processor, our Wafer-Scale Engine-3. CS-3s are quickly and easily clustered together to make the largest AI supercomputers in the world, and make placing models on the supercomputers dead simple by avoiding the complexity of distributed computing. Cerebras Inference delivers breakthrough inference speeds, empowering customers to create cutting-edge AI applications. Leading corporations, research institutions, and governments use Cerebras solutions for the development of pathbreaking proprietary models, and to train open-source models with millions of downloads. Cerebras solutions are available through the Cerebras Cloud and on-premises. For further information, visit cerebras.ai or follow us on LinkedIn, X and/or Threads.

Contacts

Cerebras Systems


Release Versions

Contacts

More News From Cerebras Systems

Cerebras Systems Announces Filing of Registration Statement for Proposed Initial Public Offering

SUNNYVALE, Calif.--(BUSINESS WIRE)--Cerebras Systems Inc. (“Cerebras”) today announced that it has filed a registration statement on Form S-1 with the U.S. Securities and Exchange Commission (“SEC”) relating to a proposed initial public offering of its Class A common stock. The number of shares of Class A common stock to be offered and the price range for the proposed offering have not yet been determined. The offering is subject to market conditions, and there can be no assurance as to whether...

Cerebras Systems Closes $850 Million Revolving Credit Facility

SUNNYVALE, Calif.--(BUSINESS WIRE)--Cerebras Systems, makers of the fastest AI infrastructure in the industry, today announced the closing of a new five-year syndicated revolving credit facility for up to $850 million. This follows the company’s $1 billion Series G financing closed in September 2025, and an additional $1 billion Series H in January 2026. “We are pleased to have closed our inaugural credit facility with the support of a syndicate of leading financial institutions,” said Bob Komi...

AWS and Cerebras Collaboration Aims to Set a New Standard for AI Inference Speed and Performance in the Cloud

SEATTLE & SUNNYVALE, Calif.--(BUSINESS WIRE)--Amazon Web Services, Inc. (AWS), an Amazon.com, Inc. company (NASDAQ: AMZN), and Cerebras Systems today announced a collaboration that will, in the coming months, deliver the fastest AI inference solutions available for generative AI applications and LLM workloads. The solution, to be deployed on Amazon Bedrock in AWS data centers, combines AWS Trainium-powered servers, Cerebras CS-3 systems, and Elastic Fabric Adapter (EFA) networking. Later this y...
Back to Newsroom