arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于可复制上下文的安全防护无法为大语言模型提供可靠的安全性

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs

Pingyu Wu, Lingyao Zhu, Weiming Zhang, Nenghai Yu

arXiv 2607.27951首次发表:更新:

发表机构

University of Science and Technology of China; Hefei AiDA Lab; Zhejiang Wanli University(中国科学技术大学; 合肥AiDA实验室; 浙江万里学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究指出基于可复制上下文的LLM安全防护存在三难困境,提出用可信凭证补充防护以预测下游使用,其结论经相关评估与程序验证具有实际意义。

AI 中文摘要

大语言模型的安全防护在看到回答将被如何使用之前就决定是否回答,这为两用任务带来了根本问题:同一个回答既可以帮助授权专业人员,也可以帮助攻击者,而攻击者可以模仿良性请求和交互历史。我们将模型释放的能力与下游使用的可用证据分开,当该证据可复制时,我们在保留有用回答的同时,得出攻击者协助的精确最坏情况下限。该结果产生了一个安全三难困境:有用的能力、可靠的安全性和开放访问无法共存。然后,我们展示了可信凭证如何通过添加难以复制的信息来补充现有安全防护,以预测实际下游使用,并确定消除该下限所需的更强条件。来自两用评估、自适应攻击和已部署可信访问程序的证据支持了这些条件的实际相关性。

英文摘要

Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and interaction history. We separate the capability released by the model from the evidence available about downstream use. When that evidence is copyable, we derive the exact worst-case floor on attacker assistance while preserving useful answers. The result yields a safety trilemma: Useful Capability, Reliable Safety, and Open Access cannot coexist. We then show how a trusted credential can complement existing safeguards by adding hard-to-copy information that predicts actual downstream use, and identify the stronger condition needed to eliminate the floor. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs supports the practical relevance of these conditions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑