ChatGPT Makes Medicine Easy to Swallow: An Exploratory Case Study on Simplified Radiology Reports
Katharina Jeblick1,2 Balthasar Schachtner1 Jakob Dexl1,3 Andreas Mittermeier1 Anna Theresa Stüber1,4 Johanna Topalis1 Tobias Weber1,3,4 Philipp Wesp1 Bastian Sabel1 Jens Ricke1 Michael Ingrisch1,3 1Department of Radiology, University Hospital, LMU Munich 2Comprehensive Pneumology Center (CPC-M), Member of the German Center for Lung Research (DZL), Munich 3Munich Center for Machine Learning (MCML) 4Department of Statistics, LMU Munich {katharina.jeblick,
Abstract
The release of ChatGPT, a language model capable of generating text that appears human-like and authentic, has gained significant attention beyond the research community. We expect that the convincing performance of ChatGPT incentivizes users to apply it to a variety of downstream tasks, including prompting the model to simplify their own medical reports. To investigate this phenomenon, we conducted an exploratory case study. In a questionnaire, we asked 15 radiologists to assess the quality of radiology reports simplified by ChatGPT. Most radiologists agreed that the simplified reports were factually correct, complete, and not potentially harmful to the patient. Nevertheless, instances of incorrect statements, missed key medical findings, and potentially harmful passages were reported. While further studies are needed, the initial insights of this study indicate a great potential in using large language models like ChatGPT to improve patient-centered care in radiology and other medical domains. ††All authors contributed equally. †† ††The title was generated with ChatGPT.
中文速览
放射科医生读不懂自己报告的患者,正在悄悄把报告发给ChatGPT让它"翻译"成大白话——这件事正在发生,但没人知道结果靠不靠谱。研究团队设计了一项探索性案例研究,让ChatGPT对三份由资深放射科医生撰写的模拟放射报告进行简化,再邀请15位放射科医生从事实准确性、完整性和潜在危害三个维度对简化结果打分评价。总体来看,大多数医生认为ChatGPT的简化报告事实基本正确、内容较为完整、对患者危害有限;但同时也发现了若干错误陈述、遗漏关键医学发现以及可能误导患者的表述。这项研究表明,以ChatGPT为代表的大语言模型(Large Language Model, LLM)在帮助患者理解医学报告、推动以患者为中心的医疗服务方面潜力可观,但其在医疗场景中的安全性和可靠性仍需更大规模的严格验证,不能盲目信任。
原文 arXiv:2212.14882;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2212.14882v1