arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18534cs.AIcs.IR

FinRCA-Bench:面向金融AI系统的证据检索与推理基准

FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

Pratik Ghawate

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出FinRCA-Bench基准,评估金融AI系统的证据检索与推理,发现检索架构影响性能,正确根本原因标签是可审计诊断的弱代理。

中文摘要 AI 辅助

大型语言模型越来越多地被用于支持金融操作,但其表面的推理性能取决于是否能获得正确的证据。在财务对账中,诊断所需的证据分布在发票、采购订单、审批、分配、付款、分类账条目和银行活动中,这些证据通过交易关系而非文本相似度关联。因此,端到端准确率可能会混淆证据获取与推理质量。我们推出FinRCA-Bench,这是一个确定性合成基准,包含2250个应付账款到银行的对账案例,涵盖14个运营表,其中包括15个因果类别中注入的1500个故障,以及750个合法或困难负例。根本原因标签和记录级证据合约对模型隐藏,允许独立于答案正确性评估检索效果。我们比较了Rules/SQL、经典机器学习、密集语义检索、确定性关系扩展和Typed Provenance Graph Retrieval(TPGR,一种仅限于持久交易关系的类型化遍历)。Rules/SQL达到84.97%的保留精确准确率,经典机器学习达到95.44%。在保持推理模型、提示和生成设置固定的情况下,仅改变检索方式,可将宏所需记录召回率从0.83%提高到77.70%,将精确16类准确率从2.05%提高到72.44%。在检索充足的情况下,结构检索失败与推理失败的比例为95:15;尽管检索不完整仍有254个正确预测,严格返回证据合约准确率仅为5.72%。在FinRCA-Bench上,检索架构强烈影响观测到的AI系统性能,而正确的根本原因标签是可审计诊断的弱代理。

英文摘要

Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase orders, approvals, allocations, payments, ledger entries, and bank activity, linked by transactional relationships rather than textual similarity. End-to-end accuracy can therefore conflate evidence access with reasoning quality. We introduce FinRCA-Bench, a deterministic synthetic benchmark of 2,250 accounts-payable-to-bank reconciliation cases spanning 14 operational tables, including 1,500 injected failures across 15 causal categories and 750 legitimate or hard-negative cases. Root-cause labels and record-level evidence contracts are hidden from the model, allowing retrieval to be evaluated independently of answer correctness. We compare Rules/SQL, classical machine learning, dense semantic retrieval, deterministic relational expansion, and Typed Provenance Graph Retrieval (TPGR), a typed traversal restricted to persisted transaction relationships. Rules/SQL reaches 84.97% held-out exact accuracy and classical ML reaches 95.44%. Holding the reasoning model, prompt, and generation settings fixed while changing only retrieval increases macro required-record recall from 0.83% to 77.70% and exact 16-class accuracy from 2.05% to 72.44%. Structural retrieval failures outnumber reasoning failures with sufficient retrieval by 95 to 15; 254 correct predictions occur despite incomplete retrieval, and strict returned-evidence contract accuracy is only 5.72%. On FinRCA-Bench, retrieval architecture strongly shapes observed AI-system performance, and a correct root-cause label is a weak proxy for an auditable diagnosis.

补充信息

↑