Evaluating Generative Models for Financial Research: Grounding, Uncertainty, and Human Review
Generative models now summarize filings, extract figures and draft research that informs investment decisions. Before that work is relied on, a research organization must answer three questions that output quality alone does not settle: whether each claim traces to a source document, whether a stated confidence tracks actual accuracy, and where a human analyst sits in the process.
This keynote covers an evaluation approach built around those three axes. Grounding, checked at claim level rather than document level, so a citation that names a real filing but does not support the sentence attached to it is caught. Uncertainty, measured against realized accuracy rather than taken at face value. And human review, defined by which outputs route to an analyst, on what trigger, and how the review is recorded.
The session covers what these measures look like in a regulated research setting, which failure modes they catch that a quality rubric does not, and how the results change what a model is permitted to do without review.
About the speaker
Dhagash Mehta
Head of Applied AI (AI, GenAI, Responsible AI) Research at BlackRock
Bio coming soon!