发表机构
ETH Zürich; Aisot Technologies Ltd(苏黎世联邦理工学院; 艾索特科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出V-FiLLM框架生成高可靠性金融推理基准,发现开源模型在金融推理深度增加及数值扰动下准确率显著下降,轻量级LoRA微调可提升模型在相关任务的表现。
AI 中文摘要
尽管现有的基准在评估大语言模型(LLMs)的STEM领域已取得实质性进展,但针对结构化数据的金融推理仍探索较少。我们提出V-FiLLM,这是一个基于真实表格的可执行计算树生成金融推理基准的框架,生成的条目答案天然正确。通过对计算树进行符号评估获取基准答案,并将其转换为自然语言问题,全程无需模型参与标注,因此可任意规模生成条目,且无标注成本,也不会继承生成器的错误率。V-FiLLM具备四个可独立控制的难度维度,包括计算深度、表达式广度、金融概念复杂度和上下文长度。通过在开源模型上评估,我们发现随着推理深度增加,准确率下降最高达51%,在对抗性数值扰动下下降最高达47个百分点,凸显了表格金融推理仍存在的挑战。我们进一步表明,在经过验证的思维链轨迹上进行轻量级LoRA微调,在保留问题上将准确率从81.1%提升至85.6%,在FinQA(Chen等人,2022a)上比基础模型提升5个百分点,表明有针对性的低成本适配是金融问答组合推理的有前景方向。
英文摘要
While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less explored. We introduce V-FiLLM, a framework that generates financial reasoning benchmarks from executable computation trees grounded in real tables, yielding items whose answers are correct by construction. Trees are evaluated symbolically to obtain ground truth and rendered into natural-language questions, removing any model from the labeling loop, so items can be generated at arbitrary scale without annotation cost and without inheriting a generator's error rate. V-FiLLM exposes four independently controllable axes of difficulty including computation depth, expression breadth, financial concept complexity, and context size. By evaluating on open-source models, we find that accuracy falls up to 51% as reasoning depth increases, and up to 47% points under adversarial numerical perturbations, highlighting remaining challenges in robust financial reasoning over tables. We further show that lightweight LoRA fine-tuning on verified chain-of-thought traces improves accuracy from 81.1% to 85.6% on held-out problems and outperforms the base model by 5% points on FinQA (Chen et al., 2022a), s), suggesting that targeted, low-cost adaptation is a promising direction for compositional reasoning in financial QA.
Comments10 pages, 6 tables, 2 figures, under review