arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型上的奖励窃取攻击

Reward Stealing Attack on Large Language Models

Jiaming Qian, Pengyang Zhou, Jiahe Xu, Chaochao Chen

arXiv 2610.06670首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出奖励窃取攻击(ReSA),利用最大熵逆强化学习从对齐模型行为中恢复安全奖励并反转生成对抗策略,跨模型泛化,显著提升攻击有效性与可迁移性。

AI 中文摘要

对大型语言模型(LLMs)的对抗性攻击旨在诱导有害内容。然而,现有方法存在计算成本高或严格模型配对依赖的问题,限制了其可扩展性和可迁移性。我们提出了奖励窃取攻击(ReSA),一种针对LLM对齐中潜在安全奖励的对抗性攻击框架。ReSA采用最大熵逆强化学习,仅从对齐模型的行为中恢复代理奖励模型。提取的奖励在推理时被反转以推导对抗性策略,通过奖励引导的解码机制高效实现。实验表明,单一恢复的奖励可跨提示和多种模型泛化,揭示根本性的对齐漏洞,使ReSA在有效性和可迁移性上显著优于现有攻击。代码可在该https URL获取。

英文摘要

Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing Attack (ReSA), an adversarial attack framework that targets the latent safety reward underlying LLM alignment. ReSA employs maximum entropy inverse reinforcement learning to recover a proxy reward model solely from the aligned model's behavior. The extracted reward is then reversed at inference time to derive an adversarial policy, efficiently implemented via a reward-guided decoding mechanism. Experiments demonstrate that a single recovered reward generalizes across prompts and diverse models to reveal a fundamental alignment vulnerability, enabling ReSA to significantly outperform existing attacks in effectiveness and transferability. The code is available at https://github.com/GarminQ/ReSA.

Comments19 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑