-

vCluster Labs Ships Stacks to Package Any AI Environment as a Managed Service

Operators define the environment once and hand it to every tenant the same way, so a new managed service is something they ship rather than staff.

SAN FRANCISCO--(BUSINESS WIRE)--vCluster Labs, the platform operators use to build and run their own AI cloud like a hyperscaler, today announced the general availability of Stacks in vCluster Platform v4.12, which automate the long list of ordered setup steps required before a managed service can go live, so an operator can deploy the same environment to every tenant.

The limit on how many managed services an operator can offer is not GPU supply. It is the setup work between an empty tenant cluster and an environment that is ready to use, stitched together by hand and repeated for every tenant.

"Every operator runs into the same wall," said Lukas Gentele, Co-Founder and CEO of vCluster Labs. "Once the cluster exists, the real project starts: registry credentials, ingress, certificates, GPU components, a control plane, and registration, all wired together by hand, in the right order, again for every tenant."

A Stack is that environment described once: which applications, in what order, what has to be healthy before the next thing starts, and which parameters the person deploying it is allowed to set. Deploy it and a tenant gets a running environment. Deploy it a hundred times and the hundredth costs what the first one did. A new managed service becomes something an operator ships rather than something they staff.

Stacks make new services faster to launch and easier to operate

A new managed service can go live in days instead of months, because the work is authoring one template rather than staffing an install for every tenant. Each additional tenant is a deployment rather than a project, which is how a GPU fleet turns into higher-margin managed products instead of raw capacity sold by the hour.

Operating them is the other half. Health, logs, upgrades, and clean tenant removal are handled by the platform, and because every tenant runs an identical environment, the team supporting the fiftieth tenant is troubleshooting the same configuration it learned on the first. Operators add new tenants without adding people to keep each environment running.

Available today

Stacks ship with ready-made examples covering the environments operators are standing up now:

  • NVIDIA Run:ai, in four certified configurations covering both a dedicated control plane for a single cluster and a shared control plane serving many tenant clusters.
  • NVIDIA Dynamo, standing up a distributed inference environment inside a tenant cluster so an operator can offer serving for large models that span many GPUs and nodes.
  • Saturn Cloud, installing the AI token factory platform inside a tenant cluster so an operator can offer per-token inference, fine-tuning, and usage billing on GPUs they already own.

"Each of these is a service an operator can put in front of customers without building the integration first," Gentele added. "That is the difference between selling capacity and selling products, and you add one without adding a team to run it. The repository is open, so the next one does not have to come from us."

"Adding a service to an AI cloud shouldn't mean starting a new integration project each time," said Hugo Shi, co-founder and CTO of Saturn Cloud. "vCluster built the Saturn Cloud Stack, so an operator can deploy it like any other service. Each customer gets private inference in an environment of their own, ready to use right away, and a small team can keep adding services without adding headcount."

Anyone can author a Stack

Stacks are not limited to integrations vCluster Labs builds. Any team whose software runs inside customer clusters can describe a correct installation once and have it land identically everywhere it is deployed. A public repository is open for exactly that, so a team that has already solved this can hand the next team a starting point.

Stacks are generally available today in vCluster Platform 4.12 with vCluster 0.37.

Supporting resources

About vCluster Labs

vCluster Labs is the platform operators use to build and run their own AI cloud, like a hyperscaler. vCluster turns raw GPUs into cluster products they can sell, from bare metal provisioning up through managed Kubernetes, Slurm and inference clusters, with tenant isolation at the infrastructure layer. The result is more margin per GPU, less dependence on a few big customers, and new managed services live in days instead of months. vCluster is trusted by fast-growing AI cloud providers including Nebius, Groq, Firmus and Corvex across 100K+ GPUs, and by enterprises including Adobe, Samsung and Deloitte.

Contacts

Media contact
Heather Fitzsimmons
Mindshare PR
heather@mindsharepr.com
650-279-4360

vCluster Labs


Release Versions

Contacts

Media contact
Heather Fitzsimmons
Mindshare PR
heather@mindsharepr.com
650-279-4360

More News From vCluster Labs

vCluster Labs Introduces vMetal to Manage Bare Metal AI Infrastructure for Neoclouds and AI Factories

SAN JOSE, Calif.--(BUSINESS WIRE)--(NVIDIA GTC Booth 206) -- vCluster Labs today introduced vMetal, a new bare metal machine management layer designed to help Neocloud providers and AI factories provision and operate GPU infrastructure at scale. vMetal automates the lifecycle of bare metal GPU servers, from initial provisioning and machine assignment to upgrades and repurposing, allowing infrastructure operators to manage physical compute with cloud-like automation. Together with vCluster’s ten...

vCluster Labs Introduces Infrastructure Tenancy Platform for AI to Maximize NVIDIA GPU Efficiency on Kubernetes Environments

ATLANTA--(BUSINESS WIRE)--(KubeCon + CloudNativeCon North America 2025, Booth #421) — vCluster Labs, the company pioneering Kubernetes virtualization, today announced its Infrastructure Tenancy Platform for AI to help organizations build and operate high-performance AI infrastructure on GPU-focused compute clusters, including support for NVIDIA DGX systems. The company’s new Reference Architecture for NVIDIA DGX systems is now available, offering architectural guidance for building secure, scal...

vCluster and Netris Announce Partnership to Deliver Full-Stack Kubernetes Multi-Tenancy for AI Infrastructure

WASHINGTON--(BUSINESS WIRE)--(NVIDIA GTC 2025 Booth #526) – vCluster Labs (formerly LoftLabs), the company pioneering Kubernetes virtualization, and Netris, the leading Network Automation, Abstraction, and Multi-Tenancy (NAAM) Platform for GPU Clouds and Enterprise AI Factories, today announced a strategic partnership to deliver the industry’s first full-stack Kubernetes multi-tenancy solution for AI infrastructure, enabling hard isolation of tenant clusters that span both physical and virtual...
Back to Newsroom