arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08934cs.CL

当模型屈从于错误答案:多项选择问答中来源属性线索的稳健性审计

When Models Defer to Wrong Answers: A Robustness Audit of Source-Attributed Cues in Multiple-Choice QA

发表机构印度理工学院 · 卡内基梅隆大学 · 亚马逊云科技AI原生部门
查看机构详情
  • Indian Institute of Technology(印度理工学院)
  • Carnegie Mellon University(卡内基梅隆大学)
  • Amazon Web Services AI Native(亚马逊云科技AI原生部门)

机构由 AI 辅助整理,请以论文原文为准。

Manikandan Ravikiran, Siddharth Vohra

首次发表
浏览论文内容

中文总结 AI 辅助

本研究审计多项选择问答中来源线索对答案稳定性的影响,提出NC-MCAR指标,发现专家模板导致41.1%的答案不稳定率,表明未经验证的来源主张可压倒基于证据的答案。

中文摘要 AI 辅助

语言模型经常接收一个问题以及关于另一个来源回答了什么的主张。我们审计了此类主张是否会破坏多项选择问答中的答案稳定性。对于每个题目,我们在误导性条件下固定一个错误选项,并改变附加于该选项的线索模板。我们引入了中性条件误导性线索采纳率(NC-MCAR),该指标仅在有效线索试验中衡量模型转向该选项的情况,即同一模型在中性提示下首先选择了正确答案的试验。这是答案不稳定性的度量,而非证明模型知道答案或所有屈从行为都是非理性的证据。我们在MMLU-Pro和IndicMMLU-Pro上评估了四个指令遵循模型,涉及英语、印地语、孟加拉语、泰米尔语和泰卢固语。在220,000个输出中,专家模板产生了41.1%的总体NC-MCAR,而多数模板为12.5%。这两种条件使用相同的错误选项和最终指令。填充词准确率远高于专家错误准确率,而正确线索提示具有较高的有效响应准确率。该审计记录了在所测试的强制选择提示下与基础依据相关的答案不稳定性:一个未经核实的来源主张可以压倒先前与任务证据一致的答案。

英文摘要

Language models often receive a question together with a claim about what another source answered. We audit whether such claims destabilize answers in multiple-choice question answering. For each item, we hold one wrong option fixed across misleading conditions and vary the cue template attached to it. We introduce \emph{neutral-conditioned misleading cue adoption rate} (NC-MCAR), which measures switches to that option only on valid cued trials where the same model first selected the gold answer under a neutral prompt. This is a measure of answer instability, not proof that the model knew the answer or that all deference is irrational. We evaluate four instruction-following models on MMLU-Pro and IndicMMLU-Pro in English, Hindi, Bengali, Tamil, and Telugu. Across 220{,}000 outputs, the expert template yields 41.1\% aggregate NC-MCAR, compared with 12.5\% for the majority template. These two conditions use the same wrong option and final instruction. Filler accuracy remains well above expert-wrong accuracy, while correct-cue prompts have high valid-response accuracy. The audit documents answer instability relevant to grounding under the tested forced-choice prompts: a bare, unverified source claim can outweigh an answer that was previously consistent with the task evidence.

补充信息

↑