arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探寻“有害弃权”:一项对AI安全基准的心理测量学审计

Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark

Christopher M. Stewart, Preston Botter, Natalie Sarabosing, Muye Zhang, Rachel Phinnemore, Shalini Ghosh, Hong Shen, Hoda Heidari

arXiv 2610.12409首次发表:更新:

发表机构

Carnegie Mellon University; Indiana University; Google(卡内基梅隆大学; 印第安纳大学; 谷歌公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对AI安全基准HELM Safety,通过心理测量学测试发现HarmBench未测量单一属性,指出安全基准的总体分数会混淆不同有害行为,主张分数需先获得单一属性解读的合理性才能用于模型比较。

AI 中文摘要

安全基准通常会为一组数据集报告一个总体分数,其中每个数据集可能针对一个或多个与安全相关的属性,因此总体分数相似的模型可能具有截然不同的属性特征。在单个属性层面比较模型更具可操作性,但通常甚至不清楚单个数据集的分数是否能隔离出任何单一属性。这类属性的一个合理候选是有害弃权(harmful refusal),即模型拒绝危险或违反政策提示的倾向。我们在HELM Safety中检验它是否构成一个单一的可测量属性。使用“属性必须先于测试存在才能被测量”的构念效度框架,我们从HELM Safety中四个可能合理针对有害弃权的数据集入手,但发现其中三个已达到饱和。我们对剩余的数据集HarmBench进行两项心理测量学测试,以确定其分数背后是否存在有害弃权这类单一属性。首先,多维项目反应理论建模强烈表明HarmBench并未测量单一属性。其次,差异项目功能分析发现,来自不同开发者、具有相同弃权能力的模型在某些项目上表现不同。这些差异在特定范围匹配(scope-specific matching)下基本消失,该模式与聚合效应一致,但不足以排除特定领域的开发者差异。从更宏观的角度看,HarmBench将不同的有害行为合并为一个分数,而HELM安全总体分数进一步将HarmBench和其他数据集的分数合并为一个单一的顶级数字。任何对数据集和项目取平均的安全分数都可能以这种方式隐藏饱和状态并混淆行为。我们认为,在使用分数比较模型之前,该分数应先获得其单一属性解读的合理性。

英文摘要

Safety benchmarks typically report one overall score for a suite of datasets, each of which may target one or more safety-related attributes, so models with similar overall scores can have very different attribute profiles. Comparing models is more tractable at the level of individual attributes, yet it is often unclear whether even a single dataset's scores isolate any single attribute. One plausible candidate for such an attribute is harmful refusal, a model's tendency to refuse dangerous or policy-violating prompts. We examine whether it constitutes a single, measurable attribute in HELM Safety. Using a construct validity framework that stipulates that an attribute must exist before a test can measure it, we start with HELM Safety's four datasets that might plausibly target harmful refusal, but find that three are saturated. We subject the remaining dataset, HarmBench, to two psychometric tests to determine if a single attribute like harmful refusal could stand behind its score. First, multidimensional item response theory modeling strongly suggests that HarmBench does not measure a singular attribute. Second, a differential item functioning analysis finds items where models from different developers with the same refusal ability score differently. These flags largely disappear under scope-specific matching, a pattern consistent with aggregation effects but not sufficient to rule out domain-specific developer differences. Zooming out, HarmBench collapses distinct harm behaviors into one score, and the overall HELM safety aggregate further collapses HarmBench and scores from other datasets into a single top-line number. Any safety score that averages over datasets and items can hide saturation and conflate behaviors this way. We argue that a score should earn its single-attribute reading before models are compared with it.

CommentsCOLM 2026, AI for Measurement Science (AIMS) Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑