基于大语言模型与检索增强生成的财经新闻自动摘要:2023年秋的早期实证研究
Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)
浏览论文内容
中文总结 AI 辅助
该2023年秋的研究构建了结合LLM与RAG的财经新闻摘要流程,测试多种模型与方法,发现Summarize Chains的Falcon-7B表现最优,RAG存在缺陷,且相关失效模式至今仍具参考价值。
中文摘要 AI 辅助
股票市场分析师与投资者每日面临一项挑战:财经新闻数量过多,而时间有限。手动阅读并综合数百篇特定公司的文章并不现实,但若遗漏关键信息,会直接影响投资决策。本项目于2023年秋在乔治华盛顿大学开展,旨在探究大语言模型(Large Language Models,LLMs)能否可靠实现这一过程的自动化。我们构建了一套流程:从News API获取新闻文章,从维基百科获取公司背景,从雅虎财经获取10家主要公司(AAPL、MSFT、GOOGL、AMZN、META、TSLA、JPM、NVDA、WMT、DIS)的股价数据。由于LLMs无法直接处理数值表格,我们开发了一个简单却有效的模板,将股票数据转换为自然语言叙述。随后,我们测试了两种摘要方法(Summarize Chains与结合FAISS的检索增强生成(Retrieval-Augmented Generation,RAG)),针对新闻采用三款开源模型(Falcon-7B-Instruct、DistilBART-CNN-12-6、BART-Large-XSum),针对股票摘要采用GPT(text-davinci-003)。结果显示,采用Summarize Chains的Falcon-7B表现最佳,能准确且连贯地覆盖所有新闻事件。检索增强生成(RAG)虽在理论上颇具前景,但当k值较大时,会导致Falcon模型出现严重重复,还会使BART-Large产生事实幻觉。两种基于LLM的方法在ROUGE-1指标上均优于简单的Lead-3基线。我们还搭建了Streamlit仪表盘用于交互式股票可视化。该研究完成于2023年秋,早于基于RAG的财经工具普及之前,我们记录的失效模式,尤其是小型模型在RAG下的幻觉问题,至今仍具参考价值。
英文摘要
Stock market analysts and investors face a daily challenge: too much financial news, too little time. Manually reading and synthesizing hundreds of company-specific articles is impractical, yet missing key information can directly affect investment decisions. This project, conducted at George Washington University in Fall 2023, explores whether Large Language Models can automate this process reliably. We built a pipeline that pulls news articles from the News API, company background from Wikipedia, and stock price data from Yahoo Finance for ten major companies (AAPL, MSFT, GOOGL, AMZN, META, TSLA, JPM, NVDA, WMT, DIS). Because LLMs cannot directly process numerical tables, we developed a simple but effective template that converts stock data into natural language narratives. We then tested two summarization approaches (Summarize Chains and Retrieval-Augmented Generation with FAISS) across three open-source models (Falcon-7B-Instruct, DistilBART-CNN-12-6, BART-Large-XSum) for news, and GPT (text-davinci-003) for stock summaries. Falcon-7B with Summarize Chains gave the best results, covering all news events accurately and coherently. RAG, while promising in theory, caused severe repetition in Falcon and hallucinated facts in BART-Large when k was large. Both LLM-based approaches outperformed a simple Lead-3 baseline on ROUGE-1. We also built a Streamlit dashboard for interactive stock visualization. The work was done in Fall 2023, before RAG-based financial tools became widespread, and the failure modes we document, particularly hallucination under RAG in smaller models, remain relevant today.
发表机构
- George Washington University(乔治·华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。