-

Rootly Acquires ThinkHive to Bring Reliability Engineering to its AI Agents

The acquisition makes Rootly's incident response agents more reliable, catching failures and drift that legacy monitoring can't

SAN FRANCISCO--(BUSINESS WIRE)--Rootly, the AI-native on-call and incident management platform trusted by companies including NVIDIA, Replit, and Canva, today announced it has acquired ThinkHive, an AI agent reliability platform. The move advances Rootly's broader goal of bringing reliability engineering to LLM workloads.

Software engineering teams spent the last decade learning to keep distributed systems reliable. They are now deploying AI agents into production, and running into a failure mode the legacy reliability playbook for deterministic systems does not handle.

ThinkHive's founders learned this the hard way. While leading AI products at Instacart, where agents served millions of customers, Nour and her team did forensic work every week just to understand why those agents were not driving better business outcomes. The failures were silent, and the bar to build agents that actually worked was high. That recurring effort is what led Nour to leave Instacart and build ThinkHive with Abdulwahab, and solve that exact problem.

ThinkHive traces every step an agent takes and evaluates whether it did its job, not just whether it returned a response. It correlates multiple signals, including metrics, traces, and evaluations, to catch the two failures that matter most in production, hallucination and drift. It clusters those failures into patterns instead of a wall of individual complaints, proposes fixes, and validates them with shadow testing before they reach a user.

For Rootly, ThinkHive serves a strategic dual purpose. First, it solves a critical customer blind spot, as companies put AI agents into production, they create new incidents that legacy monitoring simply cannot see. Second, it fortifies Rootly’s own technology. Because Rootly relies on AI during high-stakes incident response, it now leverages ThinkHive's rigorous framework, including groundedness scoring, hallucination detection, and shadow testing, to guarantee its agents are production-grade-tested before they ever touch a real incident.

That combination is one of the key pieces that sets Rootly's agentic AI apart. Every other vendor is trying to ship an agent and asking you to trust it. None can prove the agent is grounded, has not drifted, or was not quietly regressed by last week's change. Rootly can show its work.

"When an agent gives a wrong answer at scale, that is an incident, and most teams cannot see it yet," said JJ Tang, co-founder and CEO of Rootly. "ThinkHive is the team that worked out how to. We are not going to put AI into incident response and ask anyone to trust it on faith, ourselves included. That is the engineering bar we want to set for this category, and this is exactly the team we want setting it."

"Most teams think an automated eval pipeline means they have solved quality," said Nour Alkhatib, co-founder of ThinkHive. "Using AI to judge AI is like asking the same student to mark their own exam. Real reliability means tracing what actually happened and catching the failure a score hides. Rootly already brings that discipline to incident response, and already believes AI should assist the people responsible for reliability rather than replace them. That is why it is the right home for what we built."

For Rootly customers, the result is a stronger agentic AI capabilities. The same rigor ThinkHive brings to understanding agent behavior is what makes Rootly's agents better at the work that matters most during an incident: pinpointing root cause and proposing fixes a responder can act on with confidence. It deepens Rootly's proactive work too, scoring the risk of a code change against a service's incident history and live telemetry before that change ever pages someone, and determining probable incidents based on incident history and similarities. With ThinkHive's evaluation engine underneath them, Rootly's agents reason from evidence, and they do it earlier in the lifecycle.

About Rootly

Rootly is the AI-native on-call and incident management platform that helps engineering teams detect, respond to, and learn from incidents. Built to work where engineers already do, Rootly automates the manual coordination of incident response and turns every incident into a source of learning. Rootly is trusted by companies including NVIDIA, Replit, Canva, and DoorDash, and was recognized as a top-10 company on Deloitte's 2025 Technology Fast 50. Learn more at rootly.com.

About ThinkHive

ThinkHive is an AI agent reliability platform that gives teams observability and quality evaluation for the agents they run in production. Founded by Nour Alkhatib and Abdulwahab Omira, its intelligence engine correlates traces, evaluations, and business metrics so teams understand how agent behavior affects outcomes, without manual digging.

Contacts

Media contact
Adam Frank, VP of Marketing
adam@rootly.com

Rootly


Release Versions

Contacts

Media contact
Adam Frank, VP of Marketing
adam@rootly.com

Social Media Profiles
More News From Rootly

Rootly Announced as One of Deloitte's Technology Fast 50 Program Winners for 2025

TORONTO--(BUSINESS WIRE)--Rootly has been recognized by Deloitte's 2025 Technology Fast 50™ awards program with a top-10 placement for its rapid growth, entrepreneurial spirit, and bold innovation. The program recognizes Canada's 50 fastest-growing technology companies based on the highest revenue growth percentage over the past four years. Rootly CEO JJ Tang credits the company’s growth to its obsession with customer success and continued investment in AI-native features. “This recognition fro...

Rootly Announces 2025’s Top 50 People Making the World More Reliable

SAN FRANCISCO--(BUSINESS WIRE)--Rootly, the leading AI-native on-call and incident response platform, today unveiled the Reliability Top 50, an annual list recognizing the people who keep the world’s most ambitious technologies resilient at scale. When a new AI frontier model drops or an EV company pushes an autonomy update, the headlines focus on breakthroughs and valuations. But the real test is quieter: does it work when millions log in at once, when GPUs fail mid‑inference, when a global ou...

Rootly Launches AI Labs to Advance Reliability Engineering Through Community Innovation

SAN FRANCISCO--(BUSINESS WIRE)--Rootly, the leading AI-native incident response platform, today announced the launch of Rootly AI Labs, a fellow-led community dedicated to redefining the future of reliability engineering. Rootly AI Labs aims to accelerate operational excellence by developing innovative prototypes, creating open-source tools, and producing groundbreaking research. The official launch will occur at GitHub HQ in San Francisco. The event will feature keynote presentations and discu...
Back to Newsroom