arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VeriFin:用于验证大语言模型生成的财务声明的神经符号框架

VeriFin: A Neurosymbolic Framework for Verifying LLM-Generated Financial Claims

Bethel Hall, Sachi Shome, William Eiers

arXiv 2608.10213首次发表:更新:

AI 中文总结

研究针对LLM生成财务声明易出错的问题,提出神经符号框架VeriFin,基于XBRL事实与Z3验证,在XBRLFiling和FinanceBench基准上实现零错误接受,还可通过求解器反馈提升修复效果。

AI 中文摘要

大语言模型常从财务文件中生成看似合理的数值声明,但会出现报告期间、单位、行项目或公式错误的情况。验证此类声明仅重查算术运算是不够的:验证器必须锚定相关事实、确定授权计算方式,并判断候选值是否符合二者要求。我们提出VeriFin,这是一个用于数值财务问答的神经符号验证框架,它将操作数锚定在已归档的XBRL事实中,从问题、文件链接库或已记录的指标定义中推导计算方式,并使用Z3检查声明。当声明不一致时,求解器生成的不可满足核心会识别出冲突的事实、公式和候选值,从而实现针对性修复。我们在XBRLFiling(一个由28家美国公司的10-K文件构建的含600个问题的新基准)和FinanceBench上评估VeriFin。在所有验证方法共享的固定候选池中,VeriFin未接受任何错误声明,而基线方法在XBRLFiling的600个声明上接受6至92个错误声明,在FinanceBench的67个声明上接受4至21个错误声明。零错误接受的结果在多个答案生成器模型中均成立。求解器生成的反馈进一步提升了修复效果,在已正确识别的案例中实现了高达69.9%的恢复率。这些结果表明,基于源文件的符号验证可为大语言模型生成的财务声明提供可靠且可审计的接受控制。

英文摘要

Large language models often produce plausible numerical claims from financial filings while using the wrong reporting period, unit, line item, or formula. Verifying such claims requires more than rechecking arithmetic: a verifier must ground the relevant facts, establish the authorized calculation, and determine whether the candidate value follows from both. We propose VeriFin, a neurosymbolic verification framework for numerical financial question answering that grounds operands in filed XBRL facts, derives calculations from the question, filing linkbases, or documented metric definitions, and checks claims using Z3. When a claim is inconsistent, solver-derived unsatisfiable cores identify the conflicting facts, formula, and candidate value, enabling targeted repair. We evaluate VeriFin on XBRLFiling, a new 600-question benchmark constructed from 10-K filings of 28 U.S. companies, and on FinanceBench. On fixed candidate pools shared by all verification methods, VeriFin accepts none of the incorrect claims, whereas the baselines accept 6 to 92 incorrect claims on XBRLFiling's 600 claims and 4 to 21 incorrect claims on FinanceBench's 67 claims. The zero false-accept result persists across multiple answer-generator models. Solver-derived feedback further improves repair, achieving up to 69.9% recovery among true catches. These results show that source-grounded symbolic verification can provide a reliable and auditable acceptance control for LLM-generated financial claims.

Comments9 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑