|国家预印本平台
| 注册
首页|Scaling Clinical Judgment to Evaluate Medical AI

Scaling Clinical Judgment to Evaluate Medical AI

Peter G. Brodeur Byron Crowe Anthony M. Pettinato Aashna P. Shah Adrian D. Haimovich Liam G. McCoy Daniel Restrepo Jason A. Freed Ethan Goh Thomas A. Buckley Jonathan H. Chen Laura Zwaan Katherine E. Goodman Daniel J. Morgan Raja-Elie E. Abdulnour Adam Rodman Arjun K. Manrai Zahir Kanjee

✕
Arxiv_logoArxiv

Scaling Clinical Judgment to Evaluate Medical AI

Peter G. Brodeur Byron Crowe Anthony M. Pettinato Aashna P. Shah Adrian D. Haimovich Liam G. McCoy Daniel Restrepo Jason A. Freed Ethan Goh Thomas A. Buckley Jonathan H. Chen Laura Zwaan Katherine E. Goodman Daniel J. Morgan Raja-Elie E. Abdulnour Adam Rodman Arjun K. Manrai Zahir Kanjee

作者信息

Abstract

Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific questions investigated and makes it unclear whether findings would be reproduced with a different set of evaluators. To more rigorously and scalably study clinical reasoning in AI models, here we introduce PrecepTron, an LLM fine-tuned for physician-level evaluation of open-ended responses. PrecepTron was trained using low-rank adaptation (LoRA) of a 32-billion-parameter model on a small number of physician examples. We also release GRAND-ROUNDS, a new large-scale physician-annotated benchmark of 9,217 scored responses from 160 clinicians across seven studies. We show that frontier LLMs in typical "LLM-as-a-judge" approaches often disagree with physicians and with each other, but fine-tuning PrecepTron on a small number of cases enables physician-level consistent scoring across tasks. We use PrecepTron to reproduce headline findings from five influential studies assessing LLMs for clinical care in JAMA, Science, and Nature Medicine without new human grading. Using PrecepTron, we then pose new questions about how LLMs reason in medicine that would have been infeasible with human grading alone, including measuring the diagnostic accuracy of frontier LLMs when clinical cases are provided piecemeal, even token by token. Together, PrecepTron and GRAND-ROUNDS provide a foundation for reproducible, large-scale study of how LLMs reason in medicine. All code, data, and labels are made freely available for researchers.

引用本文复制引用

Peter G. Brodeur,Byron Crowe,Anthony M. Pettinato,Aashna P. Shah,Adrian D. Haimovich,Liam G. McCoy,Daniel Restrepo,Jason A. Freed,Ethan Goh,Thomas A. Buckley,Jonathan H. Chen,Laura Zwaan,Katherine E. Goodman,Daniel J. Morgan,Raja-Elie E. Abdulnour,Adam Rodman,Arjun K. Manrai,Zahir Kanjee.Scaling Clinical Judgment to Evaluate Medical AI[EB/OL].(2026-10-01)[2026-10-06].https://arxiv.org/abs/2609.12822.

学科分类

医学研究方法
首发时间: 2026-10-01
下载量:0
|
点击量:4
段落导航相关论文