Clinical NLP that clinicians will use: grounding, specialty evaluation and workflow fit

A clinical NLP system that scores well in aggregate can still go unused. Clinicians do not adopt a tool because its F1 is high. They adopt it when its output traces back to the note it came from, when it performs on their specialty rather than on the average of all specialties, and when using it costs less time than working without it.

This keynote covers building and deploying clinical NLP and generative AI inside a hospital with those three constraints in front. Grounding: every extracted fact linked to the text that produced it, so a clinician can check it in one step. Specialty evaluation: why an aggregate score hides the departments where a model fails, and what per-specialty measurement changes about what gets deployed. Workflow fit: where the output has to appear, and what happens when it appears anywhere else.

Drawn from work at the AI Center at Sheba Medical Center, the session covers what clinical teams ask for that benchmark papers do not measure, and what it takes for a model to move from a validated pilot into routine use.

About the speaker

Rachel Wities

NLP and GenAI Lead, AI Center at Sheba Medical Center

Bio coming soon!