CliniCARE-Bench: every clinical agent tested commits when the record says defer
Retrospective clinical audit is one of the most demanding real-world tasks for clinical AI: deciding what evidence a question needs, retrieving it across structured records and free-text notes, reconciling conflicting sources, applying the governing clinical standard, and citing the evidence behind every claim. CliniCARE-Bench evaluates agents on that task through 25 clinician-authored scenarios instantiated as 750 patient cases over real longitudinal MIMIC-IV data, spanning 14 medical specialties and ten reasoning capabilities. Agents work inside a governed, fully logged tool environment and return one of four verdicts: Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous. The two indeterminate classes separate a question more evidence could settle from one that stays ambiguous after a complete review, which makes principled abstention part of the scoring standard rather than a post hoc confidence threshold.
Because every retrieval, computation and citation is replayable, the benchmark scores whether the investigation behind a verdict was defensible: claim-level evidence grounding, policy citation against a fixed corpus, and a weighted rubric of required actions and prohibited shortcuts. Across 16 agentic systems built on frontier and open-weight models, four-way accuracy runs from 65.3% to 76.1%. Defect-free accuracy, which credits a verdict only when the process behind it is sound, is 4.8 to 14.8 points lower and reorders the leaderboard. For the most affected system, roughly one in five otherwise-correct verdicts rests on a prohibited shortcut. Every system evaluated commits to a definitive answer more readily than it defers on cases the record cannot settle.
The session covers how the scenarios were built, how the reference verdicts were calibrated against blinded Clinical Board review, and what the results say about where human review has to sit in a deployed audit workflow. CliniCARE-Bench is being contributed to MedHELM, with the scenario specifications and evaluation code released openly and the patient-linked artifacts available to credentialed PhysioNet users under a data use agreement.
About the speaker
Emily Xue
Head of Enterprise AI at Scale AI
Bio coming soon!