从文档到推理:面向金融数值推理的经验证合成数据流水线与语义感知微调
From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning
- Accenture Labs(埃森哲实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对金融数值推理任务,提出经验证合成数据流水线与语义感知微调方法,结合QLoRA、新型评估指标及改进损失函数,在ConvFinQA数据集上实现问答准确率显著提升。
AI中文摘要:
金融问答(QA)已成为评估大型语言模型(LLMs)在涉及表格、图表及丰富文本叙事等复杂数据格式的特定领域任务中性能的关键基准。尽管近期进展使模型能够跨模态推理并执行多步骤算术运算,但在性能一致性和评估可靠性方面仍存在局限。尤其,Exact Match(EM)等标准评估指标常未考虑单位或格式等细微差异,从而误导性能评估。本研究中,我们提出一种通过高质量合成数据生成及使用量化低秩适配(QLoRA)微调小型语言模型(SLMs)来改进金融问答系统的综合流水线。该流水线包含针对合成问答对生成的激进数据验证,以确保合成问答对的相关性与正确性。我们引入一种新型评估指标,其匹配通过算术表达式计算得出的答案而非基准答案,从而更准确地反映模型的推理能力。此外,我们提出一种改进的损失函数,将预测表达式与参考表达式通过语义相似度、新型评估指标及标准交叉熵进行对齐,以提升性能。在基准数据集ConvFinQA上的实验结果表明,使用合成数据集与所提损失函数微调后,问答准确率获得显著提升。
英文摘要:
Financial question answering (QA) has emerged as a key benchmark for evaluating the performance of Large Language Models (LLMs) on domain-specific tasks involving complex data formats such as tables, charts, and rich textual narratives. While recent advancements have enabled models to reason across modalities and perform multi-step arithmetic operations, limitations remain in performance consistency, and evaluation reliability. In particular, standard evaluation metrics like Exact Match (EM) often fail to account for minor variations such as differences in units or formats, misleading performance assessments. In this work, we propose a comprehensive pipeline for improving financial QA systems through high-quality synthetic data generation and fine-tuning of smaller language models (SLMs) using Quantized Low-Rank Adaptation (QLoRA). Our pipeline includes aggressive data validation for synthetic question answer generation to ensure the relevance and correctness of synthetic question-answer pairs. We introduce a novel evaluation metric that matches answers computed from arithmetic expressions rather than ground-truth answers; providing a more accurate reflection of model reasoning capability. Furthermore, we propose a modified loss function that aligns predicted and reference expressions using semantic similarity, our novel evaluation metric and standard cross-entropy, resulting in improved performance. Experimental results on benchmark datasets, ConvFinQA demonstrate significant gains in QA accuracy after fine-tuning using synthetic dataset and proposed loss function.