发表机构
Fudan University; Deakin University; The University of Melbourne; City University of Hong Kong(复旦大学; 迪肯大学; 墨尔本大学; 香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出HarmProfile基准数据集,收集23个前沿LLMs的超8万条有害输出,发现其有害性与多样性随模型能力增长,潜藏对齐表面下的危险知识。
AI 中文摘要
前沿大语言模型(LLMs)的安全评估在很大程度上将有害生成视为攻击结果,而非分析对象。因此,人们对模型异常行为期间产生的有害输出知之甚少,部分原因是难以获取大规模、高质量的前沿LLMs异常行为集合。为解决这一缺口,我们提出HarmProfile,这是一个以内容为核心的基准数据集,收集了不同有害类别和模型家族的模型异常行为,并将产生的有害输出分布定义为模型级风险画像。其前提是,正如可以从话语语料库中刻画语言行为一样,模型风险也可从其安全故障的内容、严重程度和变化中刻画。HarmProfile包含来自23个前沿LLMs、13个模型家族的超过80000个经证实的人工制品,分为15个有害类别和57个子类别。利用该语料库,我们发现前沿LLMs会大规模可靠地产生有害内容,但呈现出不同的风险画像;有害性和多样性均随模型能力增长,这表明前沿LLMs可能看似安全,但在对齐表面之下潜藏着日益危险的知识。我们的源代码可在this https URL获取。
英文摘要
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model misbehavior, partly because large-scale, high-quality collections of frontier-LLM misbehavior are difficult to obtain. To address this gap, we introduce HarmProfile, a content-centric benchmark dataset that collects model misbehavior across diverse harm categories and model families, and defines the resulting harmful-output distribution as a model-level risk profile. The premise is that, just as linguistic behavior can be characterized from an utterance corpus, model risk can be characterized from the content, severity, and variation of its safety failures. HarmProfile contains over 80,000 validated artifacts from 23 frontier LLMs across 13 model families, organized into 15 harm categories and 57 subcategories. Using this corpus, we find that frontier LLMs reliably produce harmful content at scale, yet exhibit distinct risk profiles; both harmfulness and diversity grow with model capability, suggesting that frontier LLMs may appear safe yet harbor increasingly dangerous knowledge beneath the alignment surface. Our source code is available at https://github.com/fresh-ma/HarmProfile .