FinGPT: Open-Source Financial Large Language Models
Hongyang Yang1, Xiao-Yang Liu2, Christina Dan Wang3 1AI4Finance Foundation; 2Columbia University; 3New York University Shanghai Corresponding author.AI4Finance Foundation: ai4finance.org
Abstract
Large language models (LLMs) have shown the potential of revolutionizing natural language processing in diverse domains, sparking great interest in finance. However, the finance domain presents unique challenges, including high temporal sensitivity, constant dynamism, and a low signal-to-noise ratio (SNR). While proprietary models like BloombergGPT have taken advantage of their unique data accumulation, such privileged access calls for an open-source alternative to democratize internet-scale financial data.
中文速览
金融领域的大语言模型一直面临数据时效性强、市场动态快、信噪比低三大难题,而现有的头部解决方案(如BloombergGPT)依赖私有数据、不对外开放,普通研究者根本用不上。为此,研究团队推出了开源框架FinGPT,以"数据为中心"的思路构建了一套从数据采集、实时清洗、轻量微调(低秩适配LoRA)到任务评测与应用落地的完整流水线,任何人都可以用它在自己的金融数据上低成本地定制专属大语言模型。在情感分析、摘要生成、数值推理等基准任务上,FinGPT展现出可与专有模型媲美的能力,并已在智能投顾等场景中完成原型演示。这项工作的意义在于打破了金融AI的数据壁垒,让学术界、从业者和开发者都能平等地参与金融大模型的研究与创新。
原文 arXiv:2306.06031;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2306.06031v2