Clinical LLM Performance Improves by >300% When Provided With High-Quality Real-World Evidence in New Precision Medicine Benchmark
Clinical LLM Performance Improves by >300% When Provided With High-Quality Real-World Evidence in New Precision Medicine Benchmark
Atropos Health releases publicly available Precision Evidence Bench, which demonstrates that LLM performance requires novel, high-quality evidence to improve performance when patient context is included in questions and rubric evaluations
PALO ALTO, Calif.--(BUSINESS WIRE)--Atropos Health, the world’s largest creator of real-world evidence (RWE) for clinical decision support, today introduced Precision Evidence Bench, a first-of-its-kind Precision Medicine Benchmark to evaluate large language models’ (LLMs’) performance against clinical questions grounded in patient context and inclusive of patient history. This Precision Benchmark evaluates how well LLMs are able to provide evidence-based answers that match the specific patient in question.
This Atropos benchmark demonstrated that it was able to outperform foundational LLMs by >300% when leveraging Scalar Evidence Content from the Atropos Alexandria Evidence Library, containing hundreds of millions of new precision Evidence-Based Findings (pEBFs or study equivalents).
The benchmark reveals that, when requiring matched context of patient history in evaluation of LLM responses, most LLMs struggle to provide direct evidence specific to a particular patient history. However, this performance is significantly improved when responses are trained on Alexandria Evidence Library content. This highlights the limitations most LLMs face in generating precision evidence-based responses to clinical questions, due in part to the overall lack of evidence in published literature. The lack of access to actual patient data for generating scaled precision evidence prohibits major commercial LLMs from delivering accurate, evidence-based recommendations.
Today, roughly 86 percent of medical decisions are made without high-quality evidence. As a result, traditional AI benchmarks evaluate LLM models on generic medical questions, completely ignoring the patient context, history, and complexity that define real-world care. To solve this precision gap, Atropos Health tested the leading general LLMs – including GPT 5.6 Sol and Astra 6 (OpenAI), Claude Opus 5 (Anthropic), as well as 3.8 Flash (Google) – with and without access to the Atropos Alexandria Evidence Library. The benchmark evaluated how well each of these LLMs responded to hundreds of real-world precision medicine questions focused on complex patient populations, treatments, comparators, and outcomes based on the Population, Intervention, Control, Outcome (PICO) framework.
When relying solely on public literature and guidelines, standard AI models returned a complete, fully cited answer, covering all PICO criteria, in 15 percent of queries. Atropos repeated the evaluation after giving the models access to 100 million pEBFs (study equivalents), and then 500 million pEBFs(study equivalents), within its Alexandria Evidence Library. The benchmark showed that model performance improved as more precision evidence became available to the models.
“The low performance from foundational models does not necessarily mean that the models were unable to understand the questions or synthesize information. Instead, they point to a more basic limitation: an AI model cannot summarize evidence that is not available to it,” said Dr. Brigham Hyde, CEO and co-founder of Atropos Health. “As most LLM deployments in healthcare now include patient context uploaded from medical history, the bar for high-quality answers must include not just published literature and guidelines, but evidence on a specific matched patient population to meet the threshold of ‘Evidence-Based Responses’.”
The results also raise questions about how medical AI is commonly evaluated. A model may perform well on questions with established answers while still struggling to provide evidence specific to a given patient. Precision medicine benchmarks can help determine whether strong test performance translates into useful answers for real patients.
“Large language models are already very capable of interpreting and communicating medical knowledge,” said Dr. Saurabh Gombar, MD, PhD, Chief Medical Officer of Atropos Health. “But many of the most important questions for clinicians and patients have never been answered by traditional research. If we want AI to support more precise healthcare decisions, we need to give them evidence to inform those decisions.”
Atropos Health has made the 209 benchmark questions and their scoring rubric freely available on Hugging Face. Researchers, healthcare organizations, and model developers will be able to use the benchmark to evaluate new models and measure progress over time. Atropos expects to release additional updates publicly on the performance of this benchmark as well as expand the precision context to other public benchmarks in the coming months.
The release also announces that Atropos is continuing to expand the Alexandria Evidence Library, now at 500 million pEBFs(Study equivalents), which is expected to reach two billion pEBFs by the end of 2026. Combining that scaled evidence with the synthesis capabilities of leading AI models will make rigorous, individualized guidance available to more clinicians and patients at the time of a healthcare decision. Doing so can help ensure every medical decision is informed by real-world evidence.
About Atropos Health
Atropos Health is the world’s largest creator of real-world evidence (RWE) for clinical decision-making. By generating real-world evidence from actual patient outcomes and data, Atropos Health is building the evidence infrastructure that clinicians, life science researchers, and AI platforms need to practice, develop, and inform medicine, grounded in what actually works. Better healthcare is only possible with evidence, and Atropos Health is working to make it available for every patient, in every healthcare decision, every day. Atropos Health does this via a suite of content and evidence-generation tools. This includes a robust network in collaboration with data partners, GENEVA OS®, the operating system across a robust network of real-world data; Alexandria®, the world’s largest evidence library; a robust data network that provides longitudinal patient data across numerous clinical specialities, an evidence agent available throughout clinical workflows, as well as a chat tool; and evidence generation tools such as ChatRWD®, that includes human-in-the-loop services to deliver studies in minutes to days. Green Button® and the Atropos Evidence™ Agent.
To learn more about Atropos Health, visit atroposhealth.com or connect through LinkedIn or X (Twitter).

