AI 中文总结
提出VERITYGATE四门模式级忠实性检查框架及配对基准,验证LLM叙述对结构化证据的遵循,发现高失败率并评估修复效果,发布代码与数据。
AI 中文摘要
流畅的LLM解释可能不遵循来自结构化系统的证据。我们提出VERITYGATE,一个四门检查器,用于检查声明的证据ID、实体、数字和声明类型。它检查一个固定的模式;它不验证散文中的每个事实。在r=0和r=1时,我们使用GPT-4o-mini、Llama-3.3-70B和Claude Sonnet 4.6测试每个设置900个实例(450个接地-未接地配对)。在这种模式级契约下,修复前,mini的80.3%声明和Sonnet的47.9%声明失败。这些是验证器拒绝率,而非散文幻觉率。一次修复通过将mini的声明存活率从19.7%提高到28.0%,Sonnet的从52.1%提高到54.3%。每个示例的已验证声明变化为mini +0.14,Llama -0.71,Sonnet -0.47,因此存活率和输出量必须一起报告。Sonnet的第二次修复通过没有明显增益。在r=1时,Gate 4覆盖了mini、Llama和Sonnet失败声明的97.0%、98.7%和100%。小型人类研究支持这些规则,但显示模式检查与正确散文之间存在差距。一个领域特定的GPT-4o评判者测试显示顺序效应,因此它只是一个有用性检查。我们发布代码和数据。
英文摘要
Fluent LLM explanations may not follow the evidence from a structured system. We present VERITYGATE, a four-gate checker for declared evidence IDs, entities, numbers, and claim types. It checks a fixed schema; it does not verify every fact in the prose. At r=0 and r=1, we test 900 instances per setting (450 grounded-ungrounded pairs) with GPT-4o-mini, Llama-3.3-70B, and Claude Sonnet 4.6. Under this schema-level contract and before repair, 80.3% of mini claims and 47.9% of Sonnet claims fail. These are verifier rejection rates, not prose-hallucination rates. One repair pass raises claim survival from 19.7% to 28.0% for mini and from 52.1% to 54.3% for Sonnet. Verified claims per example change by +0.14 for mini, -0.71 for Llama, and -0.47 for Sonnet, so survival and output volume must be reported together. A second Sonnet pass gives no clear gain. At r=1, Gate 4 covers 97.0%, 98.7%, and 100% of failing claims for mini, Llama, and Sonnet. Small human studies support the rules but show gaps between schema checks and correct prose. A domain-specific GPT-4o judge test shows an order effect, so it is only a usefulness check. We release the code and data.
Comments16 pages, 5 figures, 7 tables. Accepted at Grounding Language Models: Learning Faithfully and Efficiently (GroundLM 2026), co-located with EMNLP 2026. Code: https://github.com/sachinkg12/yukti/releases/tag/veritygate-groundlm-2026-v1.0.0 Supplementary artifact: https://doi.org/10.5281/zenodo.22668710