当缺失即证据:评估大语言模型中对完整性敏感的负向推理能力
When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models
浏览论文内容
中文总结 AI 辅助
该研究针对大语言模型的完整性敏感负向推理能力,构建CROWN-QA评估基准,发现模型存在过度闭合等问题,提示无法持续解决错误,且部分覆盖不对称性在真实文档中依然存在。
中文摘要 AI 辅助
大语言模型(LLMs)常被要求判断某事物是否不存在于记录、列表或检索上下文中。然而,只有当证据完全覆盖查询范围时,未观察到的情况才能作为给出否定答案的依据;否则,答案应保持未知,我们将这种推理称为对完整性敏感的负向推理。我们引入CROWN-QA,其包含CROWN-Synth与CROWN-Real两部分:CROWN-Synth是受控配对核心,固定问题与观察事实,仅改变查询相对覆盖范围;CROWN-Real是带有受控覆盖变体的真实文档对比集评估。在三类LLMs中,模型表现出不稳定的闭合判断与大量过度闭合,无法可靠区分合理的否定答案(认证否定)与证据不足(未知)。CROWN-Synth的主要失败是不对称性:模型常识别出隐含完整证据,却将隐含部分证据视为已覆盖查询。提示仅在过度闭合与闭合不足之间重新分配错误,而非持续解决问题。结构化证书引出法将许多错误追溯至证据覆盖范围的错误表征。CROWN-Real显示核心的部分覆盖不对称性在真实文档内容中依然存在,但其强度及过度闭合与闭合不足的平衡随模型、提示与来源而变化。
英文摘要
Large language models (LLMs) are often asked whether something is absent from a record, list, or retrieved context. Yet non-observation licenses a negative answer only when evidence completely covers the query scope; otherwise, the answer should remain unknown. We call this completeness-sensitive negative reasoning. We introduce CROWN-QA, comprising CROWN-Synth, a controlled paired core that fixes the question and observed facts while varying only query-relative coverage, and CROWN-Real, a real-document contrast-set evaluation with controlled coverage variants. Across three LLM families, models show unstable closure judgments and substantial over-closure, failing to reliably distinguish a justified negative answer (Certified-Negative) from insufficient evidence (Unknown). The dominant CROWN-Synth failure is asymmetric: models often recognize implicitly complete evidence yet treat implicitly partial evidence as query-covering. Prompting redistributes errors between over- and under-closure rather than consistently resolving them. Structured certificate elicitation traces many errors to evidence-coverage mischaracterization. CROWN-Real shows that the core partial-coverage asymmetry persists on real-document content, while its strength and the balance between over- and under-closure vary by model, prompt, and source.