结构化扰动下语言模型的稳定性与失效行为测量
Measuring Stability and Failure Behavior in Language Models Under Structured Perturbations
浏览论文内容
中文总结 AI 辅助
本研究提出分级多类别失效感知框架,在四类模型上测试七种结构化扰动,发现模型失效层级因类别而异,冲突指令与不可能前提问题是所有模型共有的弱点,此类失效无法被标准准确率检测。
中文摘要 AI 辅助
语言模型通常仅用单一准确率评分来评判,这无法揭示输入受扰动时其性能如何下降。我们提出一种分级、多类别、感知失效的框架,用于对推理模型进行压力测试。该框架沿多级严重程度阶梯对每个问题进行扰动,涵盖七个类别:保留答案的释义、输入噪声、格式、无关上下文、上下文负载、冲突指令,以及移除可答性的知识边界类别,使弃权(不执行)成为正确响应。每项测试均经过有效性筛选并标注测量得到的严重程度,每个模型则通过各级准确率、幅度加权的稳定性,以及相对于模型自身基线定义的各类别崩溃点来表征。该框架在GSM-Symbolic所用的相同100个种子问题上实例化,扩展为4473项经有效性筛选的测试,并在四个能力层级的模型上运行,揭示了聚合评分所隐藏的结构:模型失效的层级因类别而异,而非全局统一;两类压力源在所有模型中均暴露出一致的弱点,即冲突指令和基于不可能前提的问题。此外,对不可答性的识别表现不均,在信息缺失和伪造证据方面可靠,但在不可能前提方面较弱。这些失效点在标准准确率报告中无法显现。
英文摘要
Language models are usually judged by a single accuracy score, which does not reveal how their performance degrades as inputs are perturbed. We present a graded, multi-family, failure-aware framework for stress-testing reasoning models. It perturbs each problem along a multi-level severity ladder across seven families: six that preserve the answer, paraphrase, input noise, formatting, irrelevant context, context load, and conflicting instructions, and a Knowledge Boundary family that removes answerability so that refusal becomes the correct response. Every test is validity-gated and labeled by its measured severity, and each model is summarized by per-level Accuracy, a magnitude-weighted Stability, and a per-family Collapse Point defined relative to the model's own baseline. Instantiated on the same 100 seed problems used by GSM-Symbolic, expanded into 4,473 gated tests and run on four models spanning capability tiers, the framework exposes structure that an aggregate score hides: the level at which a model fails is family-specific rather than global, and two stressors expose consistent weaknesses across all models: conflicting instructions and questions built on an impossible premise. Recognition of unanswerability is otherwise uneven, reliable on missing information and fabricated evidence but weak on impossible premises. These failure points are invisible to standard accuracy reporting.
发表机构
- NeuroLoft(神经 loft(或:神经实验室))
机构由 AI 辅助整理,请以论文原文为准。