arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于混合表格和文本的金融推理的黄金引导式程序蒸馏

Gold-Guided Programmatic Distillation for Financial Reasoning over Hybrid Tables and Text

Yun Dong, Erica Zhao, Elana Chen

arXiv 2607.14709首次发表:更新:

AI 中文总结

研究针对混合表格和文本数据的金融推理问题,基于程序蒸馏开发方法,利用黄金推导指导程序合成,经执行验证转移可靠数值推理,引入迭代恢复阶段,实验表明该框架能有效训练小模型进行可靠数值推理。

AI 中文摘要

对混合表格和文本数据进行金融问答可能需要多源推理和精确数值计算。虽然大语言模型(LLMs)可以生成中间推理步骤,但自然语言推理仍容易出现算术错误,使其成为不可靠的蒸馏监督源。基于程序蒸馏,我们开发了一种方法,使用经过执行验证的Python程序,而不是自由形式的文本推理,将可靠的数值推理从大型教师模型转移到紧凑的学生模型。它利用黄金推导来指导教师端程序合成,只保留正确执行并产生黄金答案的程序,确保高质量监督。我们还引入了一个迭代恢复阶段,重新审视教师失败的示例,使学生能够恢复并将新验证的程序纳入训练。在TAT-QA上的实验表明,我们的框架对混合金融推理非常有效。我们最好的7B学生模型在测试集上达到了87.00 EM / 87.18 F1,大大超过了72B教师模型(78.46 EM)以及传统的和强大的基于LLM的基线,包括TAGOP和TAT-LLM。这些结果表明,经过执行验证的程序蒸馏为训练较小模型以执行可靠的数值推理提供了一个有效且可扩展的框架。

英文摘要

Financial question answering over hybrid tabular and textual data may require multi-source reasoning and precise numerical computation. While large language models (LLMs) can generate intermediate reasoning steps, natural-language rationales remain prone to arithmetic errors, making them an unreliable supervision source for distillation. Building on programmatic distillation, we develop an approach that transfers reliable numerical reasoning from a large teacher model to a compact student using execution-verified Python programs instead of free-form textual rationales. It leverages gold derivations to guide teacher-side program synthesis and retains only programs that execute correctly and produce the gold answer, ensuring high-quality supervision. We further introduce an iterative recovery stage that revisits teacher-failed examples, enabling the student to recover and incorporate newly verified programs into training. Experiments on TAT-QA show that our framework is highly effective for hybrid financial reasoning. Our best 7B student achieves 87.00 EM / 87.18 F1 on the test set, substantially outperforming the 72B teacher (78.46 EM) as well as traditional and strong LLM-based baselines, including TAGOP and TAT-LLM. These results demonstrate that execution-verified programmatic distillation provides an effective and extensible framework for training smaller models to perform reliable numerical reasoning.

Comments12 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑