arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

求解与验证:捕捉税务推理系统中的矛盾

Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems

Albert Sadowski, Jarosław A. Chudziak

arXiv 2609.05928首次发表:更新:

发表机构

Warsaw University of Technology(华沙理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对税务推理系统,发现模型在输入矛盾时不会弃权,通过增加一次验证调用作为矛盾门控,以最多约5%的准确率代价恢复大部分遗漏的弃权,无需训练或外部工具。

AI 中文摘要

大型语言模型现在能在法定基准测试中,对超过90%的格式良好的案例计算出正确的纳税义务,这使得它们成为直接消费此类答案的税务咨询和合规系统的候选模型。然而,真实的法律输入常常存在缺陷:必要的事实缺失,或陈述的事实相互矛盾。在干净基准上的准确性并不能说明模型在遇到缺陷输入时的行为,而直接处理缺陷输入的系统会返回一个自信的数字,却没有任何迹象表明出了问题。这引出了两个问题:当输入有缺陷时,被要求求解案例的模型是否会弃权(不执行),而当模型不弃权时,同一个模型在被要求验证输入时能否发现缺陷?我们研究了六个近期模型在SARA衍生的税务案例上,在缺失事实和矛盾事实扰动下的表现。最强的模型在事实缺失时会弃权,但会直接计算通过注入的矛盾,在63-76%的情况下返回干净输入的答案,且没有冲突信号;而当被要求验证同一输入时,它们会标记出大多数矛盾。我们将该验证调用接入一个简单的矛盾门控:当模型报告冲突时,额外的一次调用会弃权。在所有六个模型中,它恢复了大部分被遗漏的矛盾弃权,而干净准确率的代价最多约为5个百分点,且无需训练或外部工具。因此,在格式良好的输入上的高准确性是对可靠性的一种不完整度量,而求解器遗漏的检测可以通过一次自我检查廉价地恢复。

英文摘要

Large language models now compute correct tax liabilities on over 90% of well-formed cases in statutory benchmarks, which makes them candidates for the tax-advisory and compliance systems that consume such an answer directly. Real legal inputs, however, are frequently defective: required facts are missing, or stated facts contradict one another. Accuracy on clean benchmarks says nothing about how a model behaves then, and a system that computes straight through a defective input returns a confident number with no sign that anything is wrong. This raises two questions: does a model asked to solve a case abstain when the input is defective, and when it does not, can the same model catch the defect when asked instead to verify the input? We study six recent models on SARA-derived tax cases under missing-fact and contradictory-fact perturbations. The strongest models abstain when a fact is missing but compute through injected contradictions, returning the clean-input answer 63-76% of the time with no signal of the conflict; asked instead to verify the same input, they flag most of those contradictions. We wire that verification call into a simple contradiction gate: one extra call that abstains when the model reports a conflict. Across all six models it recovers most of the missed contradiction abstention at a clean-accuracy cost of at most about 5 percentage points, with no training and no external tooling. High accuracy on well-formed inputs is therefore an incomplete measure of reliability, and the detection the solver misses is cheaply recoverable with a single self-check.

CommentsAccepted for the 38th IEEE International Conference on Tools with Artificial Intelligence (ICTAI 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑