金融证据拥挤:检索增强生成中约束诱导位移的诊断与缓解
Financial Evidence Crowding: Diagnosing and Mitigating Constraint-Induced Displacement in Retrieval-Augmented Generation
浏览论文内容
中文总结 AI 辅助
针对金融RAG中冲突证据挤占有限上下文槽位的问题,提出FinDeCrowd-RAG学习式分数校正与身份门控重排序,有效恢复被挤占的证据,提升证据包含率及问答准确率。
中文摘要 AI 辅助
检索增强生成(RAG)检索候选证据,并仅将有限的高排名子集(即top-k上下文)发送给生成器。在金融问答中,段落可能匹配查询主题,但同时与其时期、细分、指标范围或表格范围冲突。我们研究了由此产生的集合级排序失败,称之为金融证据拥挤。FinDeCrowd-Stress通过匹配的兼容和不兼容候选池来隔离此失败,同时固定查询、相关证据、排名模型、候选数量和检索预算。在包含训练期间未见公司的FinDER测试分割上,不兼容池相对于同等难度的兼容池将top-10证据包含率(Recall@10)降低了0.147。这一差距表明,冲突候选消耗了有限的上下文槽位,并挤占了支持答案的证据。随后,我们引入了FinDeCrowd-RAG,一种学习到的分数校正方法,将固定相关性分数与类型化兼容性和局部词汇竞争相结合。查询级身份门控仅在预测到更好顺序时应用校正;否则,保留原始排名。在相同的受控候选上,FinDeCrowd-RAG通过恢复候选集中已存在的证据,将top-10证据包含率从0.757提高到0.902。在没有查询特定候选插入的FinDER索引上,门控重排序将该包含率从0.420提高到0.492,而top-100检索覆盖率按设计保持为0.743。使用固定生成器时,相同的排序变化提高了FinanceBench和FinQA上的答案准确性和引用召回率。这些结果将约束诱导位移确定为可测量的RAG评估目标,并表明身份门控重排序可以恢复第一阶段检索已覆盖的证据。
英文摘要
Retrieval-augmented generation (RAG) retrieves candidate evidence and sends only a limited top-ranked subset, the top-k context, to a generator. In financial question answering, passages can match a query's topic while conflicting with its period, segment, metric scope, or table scope. We study the resulting set-level ordering failure, which we call financial evidence crowding. FinDeCrowd-Stress isolates this failure through matched compatible and incompatible candidate pools while fixing the query, relevant evidence, ranking model, candidate count, and retrieval budget. On a FinDER test split containing companies unseen during training, incompatible pools reduce top-10 evidence inclusion (Recall@10) by 0.147 relative to equally difficult compatible pools. This gap shows that conflicting candidates consume limited context slots and displace answer-supporting evidence. We then introduce FinDeCrowd-RAG, a learned score correction that combines a fixed relevance score with typed compatibility and local lexical competition. A query-level identity gate applies the correction only when it predicts a better order; otherwise, it preserves the original ranking. On identical controlled candidates, FinDeCrowd-RAG raises top-10 evidence inclusion from 0.757 to 0.902 by recovering evidence already present in the candidate set. On a FinDER index built without query-specific candidate insertion, gated reranking raises this inclusion rate from 0.420 to 0.492, while top-100 retrieval coverage remains 0.743 by design. With a fixed generator, the same ordering change improves answer accuracy and citation recall on FinanceBench and FinQA. These results identify constraint-induced displacement as a measurable RAG evaluation target and show that identity-gated reranking can recover evidence already covered by first-stage retrieval.
发表机构
- ShanghaiTech University(上海科技大学)
- University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。