Customizing General-Purpose Foundation Models for Medical Report Generation
Bang Yang Affiliation: Peng Cheng Laboratory, Shenzhen 518055, China Affiliation: ADSPLAB, School of Electronic and Computer Engineering, Peking University, Shenzhen 518055, China Asif Raza Affiliation: ADSPLAB, School of Electronic and Computer Engineering, Peking University, Shenzhen 518055, China Yuexian Zou Affiliation: ADSPLAB, School of Electronic and Computer Engineering, Peking University, Shenzhen 518055, China Tong Zhang
Abstract
Medical caption prediction which can be regarded as a task of medical report generation (MRG), requires the automatic generation of coherent and accurate captions for the given medical images. However, the scarcity of labelled medical image-report pairs presents great challenges in the development of deep and large-scale neural networks capable of harnessing the potential artificial general intelligence power like large language models (LLMs). In this work, we propose customizing off-the-shelf general-purpose large-scale pre-trained models, i.e., foundation models (FMs), in computer vision and natural language processing with a specific focus on medical report generation. Specifically, following BLIP-2, a state-of-the-art vision-language pre-training approach, we introduce our encoder-decoder-based MRG model. This model utilizes a lightweight query Transformer to connect two FMs: the giant vision Transformer EVA-ViT-g and a bilingual LLM trained to align with human intentions (referred to as ChatGLM-6B). Furthermore, we conduct ablative experiments on the trainable components of the model to identify the crucial factors for effective transfer learning. Our findings demonstrate that
原文 arXiv:2306.05642;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2306.05642v1