arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

同数引用互换:压力测试 Jev 作为金融证据裁判

Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge

Chuhong Xu, Bo Su, Ziyao Chen, Ruiyang Xu, Shimeng Dai, Xinyu Qiu

arXiv 2610.08675首次发表:更新:

发表机构

Sofia University; Indiana University; University of California, San Diego; Northeastern University; Michigan State University(索菲亚大学; 印第安纳大学; 加州大学圣迭戈分校; 东北大学; 密歇根州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过同数引用互换的受控实验,评估概率性财务验证器在数字匹配固定时区分错误角色引用与有效替代引用的能力,揭示了检测角色错误与保留有效引用之间的权衡。

AI 中文摘要

财务报告在不同时期、指标和会计行之间重复数值,这使得大语言模型生成的计算结果在数值上正确的同时,可能引用了错误的财务角色。我们使用 Jev 作为 GPT-4.1-mini 计算轨迹的源支持验证器,评估概率性证据验证在数字匹配之外增加了什么。一个基于指针的带符号数字基线解释了对精确引用检查的大部分恢复。为了隔离剩余的角色识别问题,我们保持操作数和算术不变,在同数单元格之间移动引用,并保留表达等效事实的对照。这些对比揭示了既存在通过的错误角色引用,也存在被扣留的有效替代引用。显式的列标签改善了部分错误角色决策,同时也降低了对某些等效证据的支持。一个由非作者评审员标注的、基于 36 个新源页面的后续构建扩展了这项评估,并暴露了在检测角色错误和保留有效引用之间的相同权衡。其贡献是一项受控评估,该评估识别了在数字匹配保持固定时,概率性财务验证器能区分什么。对于基于大语言模型的财务助手,它使数值正确性、引用角色支持和接受结果可以分别评估。

英文摘要

Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while citing the wrong financial role. We evaluate what probabilistic evidence verification adds beyond number matching using Jev as a source-support verifier for GPT-4.1-mini calculation traces. A signed-number-at-pointer baseline explains most recovery over exact quotation checks. To isolate the remaining role-recognition problem, we hold operands and arithmetic fixed, move citations between same-number cells, and retain controls that express equivalent facts. These contrasts reveal both wrong-role citations that pass and valid alternative citations that are withheld. Explicit column labels improve selected wrong-role decisions while also lowering support for some equivalent evidence. A constructed follow-up on 36 new source pages, labeled by a non-author reviewer, extends this evaluation and exposes the same tradeoff between detecting role errors and retaining valid citations. The contribution is a controlled evaluation that identifies what a probabilistic financial verifier distinguishes when numerical matching is held fixed. For LLM-based financial assistants, it makes numerical correctness, cited-role support and acceptance outcomes separately assessable.

Commentscounterfactual citation perturbation, evidence attribution verification, financial document question answering, Jev, LLM-as-a-judge, probabilistic source verification, tabular numerical reasoning

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑