arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解决财务数据缺失危机:用于SEC 10-K提取的生成式AI流水线

Resolving the Missing Financial Data Crisis: A Generative AI Pipeline for SEC 10-K Extraction

Prisha Nair, Roee Shraga

arXiv 2609.35864首次发表:更新:

发表机构

Worcester Polytechnic Institute(伍斯特理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出生成式AI流水线,利用多种LLM从SEC 10-K中提取财务数据,解决影响超70%公司的缺失数据问题,Qwen-2.5 14B和Llama-3.3 70B分别在不同属性上取得高F1分数。

AI 中文摘要

SEC 10-K文件包含大量财务信息,但这些信息并未被一致地捕获到结构化数据集中,从而造成了一个影响超过70%的公司和半数总市值的缺失数据问题。这可能会不成比例地使定量分析偏向于规模较小的公司,这些公司可能因可用数据有限而被排除在外。传统的财务提取方法,如正则表达式(Regex)和BERT,已被广泛使用。然而,在解析复杂的SEC 10-K文件时,这些方法非常脆弱,导致文件中存在的数据丢失,因为这些方法未考虑数据属性可能位于不同章节或脚注中。本研究评估了多种大型语言模型(LLM),包括Llama-3 8B、Qwen-2.5 14B和Llama-3.3 70B,以探究各模型在从SEC 10-K文本中提取特定属性时的优缺点。提取质量通过四个结构复杂度不同的财务变量进行评估:现金及现金等价物(表格型)、短期债务(混合型)、信贷额度(叙述型)以及研发(混合型)。结果表明,虽然像Llama-3 8B这样较小的模型在复杂负面提示下性能会下降,但将参数规模与文档复杂度对齐可获得较高的零样本准确率。Qwen-2.5 14B作为表格专家表现出色,在现金提取上F1分数达到83.33%,而Llama-3.3 70B能有效处理密集的叙述性脚注,在研发提取上F1分数达到76.92%。这一可扩展框架解决了定量金融数据集中关键信息缺口,并通过更全面的SEC 10-K文件分析消除了缺失数据偏差。

英文摘要

SEC 10-K filings contain substantial financial information that is not consistently captured in structured datasets, creating a missing-data problem affecting over 70% of firms and half of total market capitalization. This can disproportionately bias quantitative analysis against smaller firms, which may be excluded due to limited available data. Traditional financial extraction methods such as Regular Expressions (Regex) and BERT, have been widely used. However, they are highly brittle when parsing complex SEC 10-K filings, which leads to data that is existent in the files being lost since these methods do not consider that a data attribute could be located in a different section or a footnote. This study evaluates several Large Language Models (LLMs), including Llama-3 8B, Qwen-2.5 14B, and Llama-3.3 70B, to figure out individual model strengths and weaknesses when extracting specific attributes from SEC 10-K text. The extraction quality was evaluated across four financial variables of varying structural complexity: Cash and Cash Equivalents (tabular), Short-Term Debt (hybrid), Credit Facilities (narrative), and Research and Development (hybrid). Results show that while smaller models like Llama-3 8B experience performance degradation under complex negative prompting, aligning parameter scale with document complexity yields high zero-shot accuracy. Qwen-2.5 14B excels as a tabular specialist with an 83.33% F1 score on Cash, whereas Llama-3.3 70B effectively navigates dense narrative footnotes, achieving a 76.92% F1 score on R&D. This scalable framework addresses critical information gaps in quantitative finance datasets and eliminates missing-data bias through a more thorough analysis of the SEC 10-K files.

CommentsPresented as a Lightning Talk at MIT URTC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑