大型语言模型(LLMs)的财务推理是否可信?针对长期财务报表的现实测试
Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements
浏览论文内容
中文总结 AI 辅助
该研究针对LLMs的财务推理可信度,构建含对抗陷阱的FinIndices基准测试,发现其存在知识与结构瓶颈,监督微调可部分恢复结构化逻辑。
中文摘要 AI 辅助
大型语言模型(LLMs)是具备真正的结构化推理能力,还是仅依赖表层模式匹配?对数值精度和长上下文多步逻辑要求严苛的金融领域,是理想的测试平台。现有基准无法捕捉现实产业的复杂性,主要依赖裁剪后表格的选择题或单跳问答,忽略了复杂的跨报表动态和时间性反累积。为填补这一空白,我们推出FinIndices——一个针对未裁剪财务报表(最长达32K tokens)评估数据处理保真度的大规模基准。通过带对抗陷阱的自动化合成流程,FinIndices涵盖单指标计算和表格指标制表,用于测试复杂领域、时间和口径推理。我们的评估揭示了LLMs的两个严重漏洞:第一,“知识瓶颈”:尽管预训练时记忆了公式,模型却表现出脆弱的模式匹配能力,移除明确公式提示会导致性能崩溃(例如Gemini-3.1-Pro在表格任务上从70.70%降至38.22%),暴露了时间性反累积和存量-流量口径不匹配的致命缺陷;第二,“结构瓶颈”:生成多指标、多周期表格的高强度认知负荷会消耗推理能力,在结构压力下,能完美执行孤立推导的LLMs会退化为浅层启发式,例如提取错误的相邻列或用惰性字面算术替代深层会计调整。最后,监督微调(SFT)带来了显著的无提示增益(单指标+8.54%,表格指标+3.82%),验证了以数据为中心的对齐可部分恢复结构化逻辑。
英文摘要
Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and multi-step logic over long contexts, is an ideal testbed. Existing benchmarks fail to capture real-world industrial complexity, predominantly relying on multiple-choice questions or single-hop QA over cropped tables while ignoring intricate cross-statement dynamics and temporal de-cumulation. To bridge this gap, we introduce FinIndices, a large-scale benchmark evaluating data-processing fidelity over uncropped financial statements (up to 32K tokens). Utilizing an automated synthesis pipeline with adversarial traps, FinIndices encompasses Single-Index computation and Table-Index tabulation to test complex domain, temporal, and caliber reasoning. Our evaluation reveals two severe LLM vulnerabilities. First, a "Knowledge Bottleneck": despite memorizing formulas during pre-training, models demonstrate fragile pattern matching. Removing explicit formula hints causes performance to collapse (e.g., Gemini-3.1-Pro drops from 70.70% to 38.22% on table tasks), exposing fatal flaws in temporal de-cumulation and stock-flow caliber mismatch. Second, a "Structural Bottleneck": the intense cognitive load of generating multi-metric, multi-period tables actively drains reasoning capacity. Under structural pressure, LLMs that flawlessly execute isolated derivations regress to shallow heuristics, such as fetching incorrect adjacent columns or substituting deep accounting adjustments with lazy literal arithmetic. Finally, Supervised Fine-Tuning (SFT) yields substantial zero-hint gains (+8.54% Single, +3.82% Table), validating that structured logic can be partially restored via data-centric alignment.