arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00519cs.CRcs.AI

防护措施起作用了吗?大语言模型系统就更安全了吗?

The Safeguard Worked. Is the LLM System Safer?

  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Pingyu Wu, Weiming Zhang, Nenghai Yu

AI总结:

该研究指出LLM防护措施的局部评估指标存在不对称性,需结合攻击者适应后的残留风险,强调防护研究应聚焦于提升部署系统的实际安全性。

AI中文摘要:

部署后的大语言模型(LLM)服务的防护措施通过拒绝率、攻击成功率和政策违规率进行评估。这些比率描述了控制措施在测试请求上的表现。而部署后的系统需要回答一个不同的问题:对于不断调整策略或找到其他入侵方式的攻击者,该服务仍能为有害任务提供多少帮助?我们确定每个报告结果对该问题的含义,从而能在同一部署标准下比较不同防护措施家族的结果。证据要求存在强烈的不对称性:一次从部署服务中获取有害帮助的攻击,就足以证明此类帮助仍然存在,且此类攻击在编码记录中反复出现。要证明残留的此类帮助很少,不能仅依靠防护措施自身的数值,还需要关于防护措施执行其局部功能后,周围系统仍允许什么的证据。在深度编码的主张中,只有少数支持或推导了此类证据,且其中一项主张对其范围内的残留情况进行了界定。因此,更好的局部分数本身并不能说明部署系统更安全。防护措施研究不能止步于提高局部分数,必须判断其改进是否真的让部署后的系统更安全。

英文摘要:

Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates. Those rates characterize how a control performed on the requests it was tested on. A deployment has to answer a different question: how much help with harmful tasks the service still gives an attacker who keeps adapting or finds another way in. We determine what each reported result implies for that question, allowing results from different safeguard families to be compared under one deployment criterion. The evidence requirements are strongly asymmetric. One attack that obtains harmful help from the deployed service suffices to establish that such help remains, and such attacks appear repeatedly in the coded record. Establishing that little remains cannot follow from the safeguard's own numbers alone; it also requires evidence about what the surrounding system still allows after the safeguard performs its local function. Such evidence is supported or derived in only a small minority of the depth-coded claims, and one such claim bounds its scoped residual. A better local score is therefore not, by itself, a stronger claim about the deployment. Safeguard research cannot stop at raising local scores; a gain has to be judged by whether it makes a deployed system any safer.

↑