蒸馏防御在强化学习后轻易失效
Distillation Defenses Easily Break After Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
本文指出蒸馏攻击后经强化学习可轻易突破现有防御,提出更现实的威胁模型,并证明简单攻击即可窃取推理能力,强调防御需考虑后续训练。
中文摘要 AI 辅助
蒸馏攻击复制了闭源大型语言模型的推理能力,使恶意行为者能够以低成本复制最先进的性能。攻击者系统地收集大量前沿模型的推理轨迹,然后在这些轨迹上训练(即“蒸馏”)自己的模型。现有的针对蒸馏攻击的防御措施通常在蒸馏后立即进行评估,隐含地假设攻击者不会进一步训练他们的模型。在本文中,我们认为更现实的威胁模型应包括蒸馏后的强化学习进一步训练。错误指定的威胁模型可能带来虚假的安全感——一些在蒸馏后看似有效的防御措施,在随后的强化学习后可能被攻破。实际上,强化学习降低了蒸馏攻击有效的门槛。我们证明,简单的攻击可以利用从当前API轻易获取的数据,从现有的闭源语言模型中窃取推理能力,其推理改进相当于提取完整隐藏轨迹的更复杂攻击。结果表明,任何泄漏足够信息以重建近似推理轨迹的蒸馏防御措施可能都是无效的。最后,我们讨论了更广泛的影响以及可能更有效的批量级蒸馏防御措施。
英文摘要
Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e., "distill") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security -- some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.
发表机构
- University of Oxford(牛津大学)
- ELLIS Institute Tübingen(ELLIS 蒂宾根研究所)
- MPI for Intelligent Systems(马克斯·普朗克智能系统研究所)
机构由 AI 辅助整理,请以论文原文为准。