Applying Large Language Models for Causal Structure Learning in Non Small Cell Lung Cancer
\NameNarmada Naik1 1dātma Health Science, Beaverton, OR, USA, 2Earle A. Chiles Research Institute, Providence Cancer Institute, Portland, OR, USA \NameAyush Khandelwal1 1dātma Health Science, Beaverton, OR, USA, 2Earle A. Chiles Research Institute, Providence Cancer Institute, Portland, OR, USA \NameMohit Joshi1 1dātma Health Science, Beaverton, OR, USA, 2Earle A. Chiles Research Institute, Providence Cancer Institute, Portland, OR, USA \NameMadhusudan Atre1 1dātma Health Science, Beaverton, OR, USA, 2Earle A. Chiles Research Institute, Providence Cancer Institute, Portland, OR, USA \NameHollis Wright1 1dātma Health Science, Beaverton, OR, USA, 2Earle A. Chiles Research Institute, Providence Cancer Institute, Portland, OR, USA \NameKavya Kannan1 1dātma Health Science, Beaverton, OR, USA, 2Earle A. Chiles Research Institute, Providence Cancer Institute, Portland, OR, USA \NameScott Hill1 1dātma Health Science, Beaverton, OR, USA, 2Earle A. Chiles Research Institute, Providence Cancer Institute, Portland, OR, USA \NameGiridhar Mamidipudi1 1dātma Health Science, Beaverton, OR, USA, 2Earle A. Chiles Research Institute, Providence Cancer Institute, Portland, OR, USA \NameGanapati Srinivasa1 1dātma Health Science, Beaverton, OR, USA, 2Earle A. Chiles Research Institute, Providence Cancer Institute, Portland, OR, USA \NameCarlo Bifulco2 1dātma Health Science, Beaverton, OR, USA, 2Earle A. Chiles Research Institute, Providence Cancer Institute, Portland, OR, USA \NameBrian Piening2 1dātma Health Science, Beaverton, OR, USA, 2Earle A. Chiles Research Institute, Providence Cancer Institute, Portland, OR, USA \NameKevin Matlock1 1dātma Health Science, Beaverton, OR, USA, 2Earle A. Chiles Research Institute, Providence Cancer Institute, Portland, OR, USA
Abstract
Causal discovery is becoming a key part in medical AI research. These methods can enhance healthcare by identifying causal links between biomarkers, demographics, treatments and outcomes. They can aid medical professionals in choosing more impactful treatments and strategies. In parallel, Large Language Models (LLMs) have shown great potential in identifying patterns and generating insights from text data. In this paper we investigate applying LLMs to the problem of determining the directionality of edges in causal discovery. Specifically, we test our approach on a deidentified set of Non Small Cell Lung Cancer(NSCLC) patients that have both electronic health record and genomic panel data. Graphs are validated using Bayesian Dirichlet estimators using tabular data. Our result shows that LLMs can accurately predict the directionality of edges in causal graphs, outperforming existing state-of-the-art methods. These findings suggests that LLMs can play a significant role in advancing causal discovery and help us better understand complex systems.
中文速览
肺癌患者的电子健康记录与基因组数据中,因果关系的方向往往难以用传统算法准确判断——这正是本文要攻克的难题。研究者以455名非小细胞肺癌(Non-Small Cell Lung Cancer, NSCLC)患者的多模态数据为实验场,将大语言模型(Large Language Model, LLM)GPT-4引入因果结构学习,让它充当领域专家,通过提示词工程逐步生成并迭代修正变量间的有向无环图(DAG)。用贝叶斯狄利克雷等效均匀(BDeu)评分对各模型与真实数据的拟合度打分后,LLM生成的最优因果图得分为-4150,而经典的NOTEARS和PC算法分别仅达到-6886和-6092,差距显著。这一结果表明,LLM能够有效弥补纯数据驱动方法缺乏先验知识的短板,在生物标志物发现和个性化治疗决策等医疗AI场景中具有重要的应用潜力。
原文 arXiv:2311.07191;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2311.07191v1