The evolution of MedHELM from single-task evaluations to multi-step workflows
MedHELM is an open source project for task-based evaluation of AI models for completing a broad range of medical tasks. Its clinician designed taxonomy of tasks spans clinical decision support, document generation, patient communication, research assistance and administration, which is accompanied by a benchmark suite of multiple public and private datasets.
Until recently, the benchmarks and scoring has been single-turn. Today we announce the integration of two new datasets, HealthAdminBench and PhysicianBench, as well as an update to the evaluation harness to extend the MedHELM framework to evaluate agents that plan, call tools and work across many turns to accomplish and end-to-end workflow. We continue the focus on the full range of tasks in healthcare spanning both administrative and patient care activities.
HealthAdminBench, comprises 135 multi-step workflows across the three domains of Prior Authorization, Appeals and Denials Management, and Durable Medical Equipment (DME) Order Processing. Each workflow is decomposed into fine-grained, verifiable subtasks, yielding 1,698 evaluations. Across seven agent configurations under multiple prompting and observation settings subtask performance reached 83% but end-to-end reliability remains low (36%).
PhysicianBench, takes 100 long-horizon tasks adapted from real primary-care-to-subspecialty consultations across 21 specialties. Tasks span 21 specialties (e.g., cardiology, endocrinology, oncology, psychiatry) and diverse workflow types (e.g., diagnosis interpretation, medication prescribing, treatment planning), requiring an average of 27 tool calls per task. Across 13 proprietary and open-source LLM agents, the best-performing model achieves only 46% success rate (pass@1), while open-source models reach at most 19%.
We will discuss how these two new benchmarks work in practice in MedHELM, where long-horizon work breaks, and how to run these new public benchmarks yourself.
About the speaker
Nigam Shah
Chief Data Scientist at Stanford Health Care
Bio coming soon!
Alin Blidisel
Engineering Technical Lead at Pacific AI
Bio coming soon!