arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17081cs.CL

EviScope:用于忠实且高效的基于事实的语言模型的配对反事实证据诊断

EviScope: Paired Counterfactual Evidence Diagnostics for Faithful and Efficient Grounded Language Models

Suryadeep Singh Deswal

首次发表
浏览论文内容

中文总结 AI 辅助

EviScope通过配对反事实证据基准,诊断基于事实的语言模型在证据变化下的忠实性,揭示答案准确性隐藏的接地缺陷,并评估显式证据门控的效能。

中文摘要 AI 辅助

基于事实的语言模型系统通常通过最终答案的准确性进行评估,然而一个正确的答案可能缺乏支持、来源于错误的资料,或在证据不足或矛盾时产生。我们引入了EviScope,一个配对反事实基准,它在保持问题不变的同时,添加、移除、干扰或反驳其证据。EviScope-v1.1包含40个四条件四重奏,带有修复后的反事实声明和用于自动评估的跨度级支持标签。在来自Qwen2.5-7B、Llama 3.1 8B和Gemini 3.5 Flash的960个金标盲生成中,配对指标揭示了答案准确性所隐藏的模型依赖的接地行为。在两个本地开放模型上,显式证据-动作门控在QCS上表现不如普通RAG:Qwen为0.15对0.50,Llama为0.10对0.375。Gemini在两种提示下达到0.944的联合成功率,但在矛盾插入后仍回答了5%的冲突案例。因此,EviScope区分了无支持的作答、冲突盲区和错误的非作答动作,而不仅仅是对答案进行评分。

英文摘要

Grounded language-model systems are often evaluated by final answer accuracy, yet a correct answer can be unsupported, drawn from the wrong source, or produced when evidence is insufficient or contradictory. We introduce EviScope, a paired counterfactual benchmark that holds the question fixed while adding, removing, distracting, or contradicting its evidence. EviScope-v1.1 contains 40 four-condition quartets with repaired counterfactual claims and span-level support labels for automatic evaluation. Across 960 gold-blind generations from Qwen2.5-7B, Llama 3.1 8B, and Gemini 3.5 Flash, paired metrics expose model-dependent grounding behavior that answer accuracy hides. On two local open models, an explicit evidence-action gate underperforms vanilla RAG on QCS: 0.15 vs. 0.50 for Qwen and 0.10 vs. 0.375 for Llama. Gemini reaches 0.944 joint success under both prompts, yet still answers 5% of conflict cases after contradiction insertion. EviScope therefore distinguishes unsupported answering, conflict blindness, and wrong non-answer actions rather than scoring answers alone.

发表机构

  • Indian Institute of Technology Roorkee(印度理工学院鲁尔基分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑