面向金融推理的数据中心后训练:挖掘、蒸馏与可验证学习
Data-Centric Post-Training for Financial Reasoning: Mining, Distillation, and Verifiable Learning
AI总结:
本文提出一种数据中心后训练流水线,通过挖掘、蒸馏和知识图谱生成构建金融推理语料,并采用自蒸馏SFT和模型合并等保留感知方法,在FINESSE-Bench上提升金融推理性能,避免普通SFT的退化。
AI中文摘要:
金融文本、教科书和问答对数量丰富,但其中只有一小部分可直接用于以推理为中心的后训练。现有的问答对往往缺乏显式推理、充分的上下文或可靠可验证的答案,而教科书则必须先转化为合成的训练示例。我们提出了一种数据中心流水线,通过挖掘开源推理轨迹、蒸馏金融指令数据以及从金融教育材料中生成知识图谱引导的问答对,来构建互补语料库。在语义去重之后,三个轻量级序列分类器选择与金融相关的示例,拒绝不明确的问题,并识别适合使用紧凑的基于规则的验证器进行强化学习的任务。对于模型适配,我们研究了监督微调和强化学习,同时使用自蒸馏微调和后训练模型合并来防止起始模型中已有的金融能力丢失。我们使用FINESSE-Bench评估适配后的语言模型,报告聚合性能及其相对于起始检查点的变化。在所选比较中,普通SFT使FINESSE-Bench准确率降低3.2-4.0个百分点,而自蒸馏SFT相对于相应的起始模型提高了1.0-2.8个百分点。等权合并相对于其SFT父模型恢复了3.0个百分点,并最终比原始模型高出0.9个百分点;在困难任务上的GRPO在自蒸馏SFT之后增加了0.4个百分点,或直接应用于可验证任务时增加了3.0个百分点。这些结果表明,保留感知的适配可以在不出现普通SFT后观察到的退化的情况下改善金融推理。
英文摘要:
Financial text, textbooks, and question-answer pairs are abundant, but only a small fraction is directly usable for reasoning-focused post-training. Existing QA pairs often lack explicit reasoning, sufficient context, or reliably verifiable answers, while textbooks must first be transformed into synthetic training examples. We present a data-centric pipeline that constructs complementary corpora by mining open-source reasoning traces, distilling financial instruction data, and generating knowledge-graph-guided question-answer pairs from financial educational material. After semantic deduplication, three lightweight sequence classifiers select finance-relevant examples, reject under-specified questions, and identify tasks suitable for reinforcement learning with compact rule-based verifiers. For model adaptation, we study supervised fine-tuning and reinforcement learning, while self-distilled fine-tuning and post-training model merging are used to prevent the loss of financial capabilities already present in the starting model. We evaluate the adapted language models using FINESSE-Bench, reporting aggregate performance and changes relative to their starting checkpoints. Across the selected comparisons, ordinary SFT reduces FINESSE-Bench accuracy by 3.2-4.0 percentage points, whereas self-distilled SFT improves over the corresponding starting models by 1.0-2.8 points. Equal-weight merging recovers 3.0 points over its SFT parent and finishes 0.9 points above the original model; GRPO on hard tasks adds 0.4 points after self-distilled SFT or 3.0 points when applied directly to verifiable tasks. These results show that retention-aware adaptation can improve financial reasoning without the regressions observed after ordinary SFT.