Molt AI Launches Fisher to Test What AI Agents Do Under Pressure
Molt AI Launches Fisher to Test What AI Agents Do Under Pressure
New technical white paper documents more than 50,000 multi-turn adversarial episodes and shows how final-response grading can miss agent actions that have already occurred.
MIAMI--(BUSINESS WIRE)--Molt AI Corp. ("Molt"), an enterprise AI assurance company, today launched Fisher, its independent Agent Assurance Assessment for tool-using AI agents, and published the technical white paper documenting the evidence and methodology behind it. Fisher uses adaptive, multi-turn adversarial testing to examine agent behavior at the level where consequences occur — tool calls, retrievals, writes, API activity, and boundary crossings — where those signals are observable. Molt also announced the close of $1 million in pre-seed financing to expand Fisher and its agent-assurance research and delivery capacity.
Enterprise AI agents no longer just generate text. They call tools, read private data, query databases, send messages, and change systems in ways that persist after the conversation ends. Traditional model evaluations often emphasize final outputs; tool-using agents require evaluation of the full trajectory, including the data they retrieve, the tools they call, and the changes they make. An agent can produce an appropriate final response after it has already queried protected data or taken an action outside its intended boundary. An evaluator reading only the transcript scores that session as a pass.
"A refusal can hide a completed action," said Greg Frank, co-founder of Molt AI and creator of Fisher. "If an agent has already queried a protected row or sent a message, its final sentence is not the security outcome. Fisher follows the action evidence, replays the path, and states exactly what the record supports."
What Fisher does differently
Fisher runs adaptive, multi-turn adversarial campaigns against an organization's agents and evaluates the strongest evidence available. Where action evidence is observable, it grades tool calls, retrievals, writes, API activity, and boundary crossings. Where an endpoint is opaque, Fisher claims only what the available response evidence or owner-controlled synthetic evidence supports.
Evidence is labeled at three levels — exploratory, observed, and confirmed — and a candidate weakness becomes a confirmed finding only when it reproduces under recorded replay conditions. For confirmed findings, Fisher proposes configuration-level remediation where the target supports one, then re-attacks the fix, re-running the original attack and, where authorized, a fresh adaptive one, and reports the outcome as fixed, partial, bypassed, or unknown. The report records what ran and what remains unknown.
Fisher also retains reusable units of attack intent with their provenance, so authorized strategy evidence from prior campaigns can help later campaigns start smarter. Strategy variants gain weight from observed outcomes, not novelty. Customer-specific transcripts, findings, and artifacts remain governed by configured workspace, retention, and sharing rules.
What the white paper found
As of July 2026, Fisher has run more than 50,000 multi-turn adversarial episodes and 500,000+ conversation turns across more than 20 model architectures, with hundreds of retained attack strategies and variants. In matched experiments that held the target model, scenario, scoring rule, and judge constant, adaptive tactics and refusal pivots increased verified attack success by 6.6 to 14.6 percentage points across two model families and two scenarios.
In a controlled historical comparison on the same database agent, a widely used open-source evaluation platform graded several sessions as defended because the final text refused the request, while the action trace showed password and API-key records had already been queried. Fisher caught the action-level failures. That result belongs to its fixed target, test set, and date: it is evidence of a grading blind spot, not a universal performance claim about either product. The methodology and complete results are published at https://moltaicorp.com/paper.
"The moment an agent can act on a company's systems, testing it with adversarial prompt libraries or merely reading its final answer stops being enough," said Walton Comer, Co-Founder and CEO of Molt AI. "You need to know what the agent actually did, what it could do in the future, whether you can make it fail the same way again, and whether the fix still holds when it's attacked in a new way. That's a different standard than reviewing transcripts, and it's the standard enterprises are going to be held to by auditors and compliance frameworks."
"In thirty years of testing corporate defenses, nothing has moved a room like watching their own systems fail in front of them," said Gideon Lenkey, President and Co-Founder of Ra Security and a past president of InfraGard's New Jersey chapter, who has been contracted by the FBI to provide advanced training. "That's what Fisher does for agents: follow the actions, reproduce the failure, then attack the fix and see whether it holds. Adoption has gotten out ahead of controls at nearly every client I have. Fisher is how we close that gap with evidence instead of argument."
Comer is a technology entrepreneur and former quantitative trader. He co-founded Lucid, the research-technology company acquired by Cint in 2021 for approximately $1.07 billion, and XBTO, a digital assets firm where he served as Chief Investment Officer, and was an early investor and advisor in Deribit, which Coinbase acquired in a transaction valued at approximately $2.9 billion. He is a co-founder and the Chairman of Roundtable (RTB), and he began his career as a quantitative analyst at SAC Capital, a hedge fund. Molt AI's founding team includes senior AI engineers from Nvidia, product managers from Microsoft, and published AI/ML researchers.
Molt's $1 million pre-seed financing is closed and was provided by experienced founders and operators of technology firms, long-term supporters, and individual investors. Participants include Patrick Comer, founder of Lucid and chief executive officer of Cint, which acquired Lucid in 2021. Proceeds will support expansion of Fisher, enterprise assessment and integration capacity, and research into agent behavior, action-level evaluation, and verified remediation. Fisher's Prove and Remediate capabilities are available today; Discover is in preview with design partners.
About Molt AI
Molt AI is an enterprise AI assurance company. Fisher, its independent Agent Assurance Assessment, tests tool-using AI agents under adversarial conditions, documents evidence-backed findings, and verifies supported remediations through retesting. Molt helps security, engineering, and risk teams establish how agents behave when given access to consequential enterprise systems. Molt is headquartered in Miami, with offices in New York. Learn more at moltaicorp.com.
Contacts
Media contact:
Rebekah Keida
Molt AI Corp
(954) 937-9731
rebekah@moltaicorp.com
