发表机构
School of Cyberspace Security, Beijing University of Posts and Telecommunications(北京邮电大学网络空间安全学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FUSE是评估大语言模型危险能力的模块化框架,通过K、D、H三个维度评估12款商用LLM,发现危险能力未随模型更新单调下降,且不同模型家族的危险能力特征存在显著差异。
AI 中文摘要
碎片化的安全评估会损害对危险AI能力的治理。我们提出了一个模块化框架,在统一协议下通过三个正交流水线——知识(K)、防御(D)和危害(H)——评估每个模型,将结果汇总为标准化的危险能力轮廓φ。可插拔模块提供场景种子、知识库、危害查询和评判标准,核心评估引擎在各领域保持不变;网络试点补充了CB评估,证明了协议的可迁移性。通过化学生物(CB)模块实例化该框架,我们评估了四个家族的12个商用大语言模型。第一项贡献是跨模型和模型家族的危险能力横向比较:三个维度呈现出截然不同的轮廓——知识水平相当的模型在拒绝韧性上存在差异,防御能力强的模型在执行合规操作时生成的有害内容并未减少;家族层面的模式进一步区分了Claude、DeepSeek和GPT模型。第二项贡献是能力演化的时间分析:对照模型发布日期追踪K、D、H,发现危险能力并未单调下降;更新的模型深化了知识,仅在防御上有部分提升,表明缩放和对齐进展并未均匀转化为安全性。跨评判者一致性(bootstrap ρ>0.79,5名评判者中的4名)和流水线正交性(K-D-H的相关性ρ∈[0.32,0.52])确立了可靠性。
英文摘要
Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular framework that evaluates each model through three orthogonal pipelines---Knowledge ($K$), Defense ($D$), and Harm ($H$)---under a unified protocol, aggregating results into a standardized dangerous-capability profile $ϕ$. Pluggable modules supply scenario seeds, knowledge banks, hazard queries, and judge rubrics, while the core evaluation engine remains unchanged across domains; the CB evaluation is complemented by a cyber pilot demonstrating protocol transfer. Instantiating the framework with a chemical-biological (CB) module, we evaluate 12 commercial LLMs from four families. Our first contribution is a horizontal comparison of dangerous capability across models and model families: the three dimensions expose sharply divergent profiles---models with comparable knowledge differ in refusal resilience, and strong defenders do not generate less harmful content when they do comply---while family-level patterns further separate Claude, DeepSeek, and GPT models. The second is a temporal analysis of capability evolution: tracking $K$, $D$, and $H$ against model release dates reveals that dangerous capability has not monotonically declined; newer models deepen knowledge while only partially improving defense, showing that scaling and alignment progress do not uniformly translate into safety. Reliability is established via cross-judge consistency (bootstrap $ρ> 0.79$, 4 of 5 judges) and pipeline orthogonality ($K$--$D$--$H$ inter-correlations $ρ\in [0.32, 0.52]$).