发表机构
Northeastern University; Georgia Institute of Technology; University of Southern California(东北大学; 佐治亚理工学院; 南加利福尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现仅用单一规范提示评估大语言模型安全会低估其不安全合规性,不同表面形式下的不安全结果并集会超出最糟单一形式,且存在模型差异,同时发布了相关数据集、代码与响应标签。
AI 中文摘要
基准分数是一种测量工具,但大多数基准仅以单一规范表面形式读取每个条目。我们询问这种读取是否准确:当条目的意图保持固定,仅其保留意义的表面形式发生变化时,规范形式的分数能否很好地估计模型行为,以及这种变化中有多少是解码/评判噪声而非有效信号?我们在安全这一高风险场景中实例化该问题,该场景没有可用于平均的黄金标签。为避免先前的混淆,我们预先制定了改写形式(无拒绝、主要非大语言模型:机器回译和Matrix-Language-Frame代码切换生成器),使相同的表面形式能到达每个模型;我们使用一个以人类为基准、与供应商无关的评判器(Claude,在不安全合规性上与人类的kappa=0.86,跨语言稳定,经GPT-4o交叉验证)对所有响应打分,并验证意图保留。在370个种子×5种表面形式×5个模型的设置中,没有任何单一转换是始终最危险的(20项每项转换的McNemar检验中,6项经校正后仍显著,多数具有保护性)。然而仅评估规范提示会低估不安全合规性:不同形式下不安全结果的并集甚至比最糟糕的单一形式高出3.3-12.9个百分点,5个模型的bootstrap 95%置信区间均不包含零,且5%-13%在规范形式下安全的种子在某些改写形式下不安全——这高于零随机性基准(在温度为0时对规范形式重新采样5次,未产生新的不安全暴露)。该差距的大小取决于模型(在Gemini 2.5 Pro上最大)。一种形式仅能恢复模型观察到的不安全表面的约53%,约三种形式可达到85%——这是该形式集的冗余性表征,而非针对定义明确的总体。良性对照(XSTest)表明这种不稳定性是双向的,尽管良性和有害的条目池未匹配。我们发布了数据集、代码和每个响应的标签。
英文摘要
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving surface form varies, does the canonical-form score estimate model behavior well, and how much of any variation is decoding/judge noise rather than signal? We instantiate this in safety, a high-stakes setting with no gold label to average toward. To avoid prior confounds, we pre-author the reformulations (refusal-free, mostly non-LLM: machine back-translation and a Matrix-Language-Frame code-switch generator) so an identical surface form reaches every model, score all responses with one human-anchored, vendor-neutral judge (Claude, kappa = 0.86 vs. human on unsafe compliance, stable across languages, cross-checked by GPT-4o), and verify intent preservation. On 370 seeds x 5 surface forms x 5 models, no single transformation is uniformly most dangerous (6 of 20 per-transformation McNemar tests survive correction, most protective). Yet evaluating only the canonical prompt underestimates unsafe compliance: the union of unsafe outcomes across forms exceeds even the worst single form by 3.3-12.9 pp, with bootstrap 95% CIs excluding zero for all five models, and 5-13% of seeds safe on canonical are unsafe under some reformulation -- above a zero stochasticity floor (canonical resampled five times at temperature 0 gives 0/370 new exposures). The size of this gap is model-dependent (largest on Gemini 2.5 Pro). One form recovers only ~53% of a model's observed unsafe surface and about three reach 85% -- a redundancy characterization of this form set, not of a defined population. A benign control (XSTest) suggests the instability is bidirectional, though the benign and harmful pools are not item-matched. We release the dataset, code, and per-response labels.
CommentsAccepted at the Sci-FM Workshop @ COLM 2026 (non-archival). Workshop version with reviews: https://openreview.net/forum?id=mZh0MqpOOC