-

Cirrascale Cloud Services Launches Production Release of the Cirrascale Inference Platform, Delivering a Complete Enterprise AI Inference Stack

New release pairs turnkey enterprise AI capabilities and governance with support for model routing, fine-tuning, and accelerator selection across NVIDIA, AMD, Qualcomm, and Tenstorrent hardware

SANTA CLARA, Calif.--(BUSINESS WIRE)--AI Infra Summit – Cirrascale Cloud Services, the expert neocloud built for Private AI, today announced the production release of the Cirrascale Inference Platform, a complete software stack for enterprise-grade AI inference. From a single serverless platform, enterprises can run open-source models, their own private models, and closed model ecosystems, including Google Gemini delivered on premises through Google Distributed Cloud and operated by Cirrascale.

“Cirrascale Inference platform gives customers working applications, spend controls, and agent guardrails on day one, running privately on the accelerator that makes the most sense for their workload.” - Alex Nataros, CTO at Cirrascale Cloud Services

Share

The platform is built for the workloads enterprises are deploying today: search, chatbots and copilots, agentic AI services, coding assistants, document intelligence, and video generation. Every pipeline is managed from a single web console and connects securely where existing data and services exist, in on-premises and hyperscaler environments.

Its web front application, inference AI capabilities, and security layer remove the months of integration work that usually stand between raw infrastructure and a tool employees can use. Enterprises get a turnkey private chat experience connected to their own knowledge base, built-in controls to manage AI spend across teams, and governance guardrails for agentic workloads, all aligned with HIPAA, SOC 2, and FedRAMP requirements where required.

“Enterprises do not struggle to stand up a model endpoint anymore. They struggle with everything around it: governance, cost control, and getting a secure application in front of employees,” said Alex Nataros, CTO at Cirrascale Cloud Services. “This release closes that gap. Customers get working applications, spend controls, and agent guardrails on day one, running privately on the accelerator that makes the most sense for their workload.”

The optimization layer of the Cirrascale Inference Platform is built to get the most out of every accelerator. It maximizes throughput, keeps latency stable as demand spikes, and lets customers run larger models on existing hardware. The result is more tokens per GPU dollar and predictable cost per token across every workload.

Additionally, the model and hardware selection layer ends the tradeoff between model choice and infrastructure commitment. The platform automatically routes each request to the right model and runs it on the best available accelerator, whether that be NVIDIA, AMD, Tenstorrent, or Qualcomm, with no code changes required to switch, giving organizations flexibility to deploy leading open models on the most appropriate hardware. Teams can also fine-tune models on their own private data without that data ever leaving their environment.

That combination sets Cirrascale apart from both hyperscalers and other GPU clouds. No other provider pairs automated, multi-vendor accelerator selection with Private AI delivery of closed models like Gemini and data centers connected close to where customers already operate.

“Hyperscalers give you their models on their hardware. Most GPU clouds give you one vendor’s silicon and leave the software to you,” said Dave Driggers, CEO and co-founder of Cirrascale Cloud Services. “We built the Cirrascale Inference Platform so enterprises never have to make that choice. They pick the model and we put it on the best hardware for the job, in a private environment, at a price their CFO can plan around.”

The Cirrascale Inference Platform is available now across Cirrascale’s U.S. and international regions. To learn more or request a demonstration, visit https://inference.cirrascale.com.

About Cirrascale Cloud Services

Cirrascale Cloud Services is the expert neocloud delivering cloud and managed services dedicated to providing tailored, state-of-the-art compute resources and high-speed storage solutions at scale. Our AI Innovation Cloud and Inference Platform services are purpose-built to enable clients to scale their training, fine-tuning, and inferencing workloads for Private AI, generative AI, large language models, and high-performance computing. To learn more about Cirrascale Cloud Services and its unique cloud offerings, please visit https://cirrascale.com.

©2026 Cirrascale Cloud Services LLC. All rights reserved. Cirrascale and the Cirrascale logo are trademarks of Cirrascale Cloud Services LLC.

Contacts

Contact Information:
pr@zmcommunications.com

Cirrascale Cloud Services


Release Versions

Contacts

Contact Information:
pr@zmcommunications.com

More News From Cirrascale Cloud Services

Cirrascale Advances Open AI Infrastructure with AMD Helios Rackscale Solution and AMD Instinct™ MI400 Series GPUs

SAN DIEGO--(BUSINESS WIRE)--Cirrascale Cloud Services, the expert neocloud built for Private AI, today announced support for the AMD Helios rackscale solution and AMD Instinct™ MI400 Series GPUs across its AI innovation cloud. The addition brings AMD rack-scale AI infrastructure, powered by AMD Instinct™ MI455X GPUs, to Cirrascale customers running large-scale inference, frontier-model training, and fine-tuning, and it extends Cirrascale's open, multi-vendor approach to accelerated computing. A...

Cirrascale Cloud Services Adds Tenstorrent Galaxy Blackhole to Its AI Innovation Cloud

SAN DIEGO--(BUSINESS WIRE)--Cirrascale Cloud Services, the expert neocloud built for Private AI, today announced the addition of Tenstorrent Galaxy Blackhole servers to its AI Innovation Cloud. The integration marks Tenstorrent's entry into broad commercial deployment and expands Cirrascale's portfolio with hardware built specifically for the demands of real-world AI. Tenstorrent Galaxy is engineered around the metrics that matter most in production AI environments: cost per token, latency per...

Cirrascale Expands Model Offerings to Include Gemini on Google Distributed Cloud with the Cirrascale Inference Platform

SAN DIEGO--(BUSINESS WIRE)--Cirrascale Cloud Services, the expert neocloud built for Private AI, today announced it is broadening its collaboration with Google Cloud to deliver Google’s Gemini models on-premises via Google Distributed Cloud, now available as part of the Cirrascale Inference Platform. Public sector organizations and enterprises can run Gemini models directly within their own infrastructure in Cirrascale data centers, keeping sensitive data behind their firewall while accessing t...
Back to Newsroom