Large Language Models Encode Clinical Knowledge
Karan Singhal Google Research, Shekoofeh Azizi Google Research, Tao Tu Google Research, S. Sara Mahdavi Google Research, Jason Wei Google Research, Hyung Won Chung Google Research, Nathan Scales Google Research, Ajay Tanwani Google Research, Heather Cole-Lewis Google Research, Stephen Pfohl Google Research, Perry Payne Google Research, Martin Seneviratne Google Research, Paul Gamble Google Research, Chris Kelly Google Research, Nathaneal Schärli Google Research, Aakanksha Chowdhery Google Research, Philip Mansfield Google Research, Blaise Agüera y Arcas Google Research, Dale Webster Google Research, Greg S. Corrado Google Research, Yossi Matias Google Research, Katherine Chou Google Research, Juraj Gottweis Google Research, Nenad Tomasev DeepMind Yun Liu Google Research, Alvin Rajkomar Google Research, Joelle Barral Google Research, Christopher Semturs Google Research, Alan Karthikesalingam Google Research, Vivek Natarajan Google Research,
Abstract
Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but the quality bar for medical and clinical applications is high. Today, attempts to assess models’ clinical knowledge typically rely on automated evaluations on limited benchmarks. There is no standard to evaluate model predictions and reasoning across a breadth of tasks. To address this, we present MultiMedQA, a benchmark combining six existing open question answering datasets spanning professional medical exams, research, and consumer queries; and HealthSearchQA, a new free-response dataset of medical questions searched online. We propose a framework for human evaluation of model answers along multiple axes including factuality, precision, possible harm, and bias.
原文 arXiv:2212.13138;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2212.13138v1