arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

稳健性的错觉:聚合准确率掩盖了与任务无关的上下文下的预测翻转

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

Yanzhe Zhang, Sanmi Koyejo, Diyi Yang

arXiv 2607.12963首次发表:更新:

发表机构

Georgia Tech; Stanford University(佐治亚理工学院; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型在含无关上下文环境中的表现,发现聚合准确率掩盖了单个示例预测的不稳定性,如随机伪词会改变部分预测,且此不稳定性受多种因素调节,揭示了尾部风险,推动对模型进行单个示例可靠性评估。

AI 中文摘要

随着大语言模型能力增强,它们越来越多地部署在上下文丰富的环境中,任务输入常伴有冗长且部分无关的上下文。在可控环境中,我们发现最先进的模型在聚合层面通常对与任务无关的上下文表现出稳健性:在基准问题前添加该上下文对整体准确率影响不大。然而,这种聚合稳定性掩盖了单个示例的显著不稳定性。即使是随机组合字符形成的语义无意义的伪词,也能在一小部分示例上显著改变模型预测,在一些示例上降低性能,在另一些上提高性能。这种双面效应在广泛的模型和数据集上持续存在,且受影响的示例很大程度上因模型而异。我们进一步表明,这种不稳定性受上下文类型、长度、测试时计算和模型开发阶段的调节。我们的发现揭示了聚合准确率掩盖下的上下文诱导的尾部风险,促使对语言模型进行单个示例的可靠性评估。

英文摘要

As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy. This aggregate stability, however, masks significant per-example instability. Even semantically meaningless pseudo-words, formed by randomly combining characters, can markedly shift model predictions on a small fraction of examples, degrading performance on some while improving it on others. This two-sided effect holds consistently across a wide range of models and datasets, yet the affected examples are largely model-specific. We further show that this instability is modulated by context type, context length, test-time compute, and model development stage. Together, our findings reveal context-induced tail risks concealed by aggregate accuracy, motivating per-example reliability evaluation of language models.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑