arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向AI安全的项目反应理论

Item Response Theory for AI Safety

Joshua Fonseca Rivera, Neil Shah, David Demitri Africa, Konstantinos Voudouris

arXiv 2608.05086首次发表:更新:

发表机构

Independent(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究将项目反应理论(IRT)应用于192个语言模型的8个安全基准,识别出三个关键因素,证明IRT可降低评估成本并审计模型,建议前沿实验室采用。

AI 中文摘要

语言模型的安全行为表现存在差异,这类差异可通过安全基准进行衡量,但聚合基准分数难以信任和解释,因为基准之间存在重复、高度相关,且模型在检测到评估时可能会弃权(不执行)。为解决这些问题,我们采用项目反应理论(IRT),这是一种从具有推断心理测量属性的项目表现中测量潜在特质的统计工具包。我们将IRT模型应用于192个语言模型的8个安全基准,开展了迄今为止最大规模的LLM安全评估心理测量分析,并得出三项结果:第一,我们发现拒绝严格程度、真实性和情境危害这三个可解释因素,解释了模型在基准间的大部分差异;第二,经心理测量选择的项目,其恢复完整基准分数的误差低于相同规模的随机子集,且几个基准仅需约10个自适应选择的项目,可将评估成本降低97%-99%;第三,IRT支持对单个模型进行审计,可用于检测简单的弃权(不执行)行为和API背后的模型变更。总体而言,我们表明IRT是一种现成的工具包,可用于解读、简化和审计安全基准,建议前沿实验室和评估人员采用。

英文摘要

Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. We fit IRT models to eight safety benchmarks across 192 language models, the largest psychometric analysis of LLM safety evaluations to date, and contribute three results. First, we find that three interpretable factors of refusal strictness, truthfulness, and contextual harm explain most of the variance between models across benchmarks. Second, psychometrically selected items recover full benchmark scores with lower error than random subsets of the same size, and roughly ten adaptively chosen items suffice for several individual benchmarks, cutting evaluation cost by 97-99%. Third, IRT supports audits of individual models, showing that it can be used to detect naive sandbagging and changes of model behind APIs. Overall, we show IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which we recommend frontier labs and evaluators adopt.

Comments15 pages, 9 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑