arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Red-TTT:用于自动越狱大型语言模型的测试时训练

Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models

Tongyan Hu, Hao Li, Xiaogeng Liu, Ruida Wang, Zhengyu Liu, Shuyao Xu, Ning Zhang, Ziyang Li, Yinzhi Cao, Bryan Hooi, Chaowei Xiao

arXiv 2610.05282首次发表:更新:

发表机构

Johns Hopkins University; National University of Singapore; Washington University in St. Louis; University of Illinois Urbana-Champaign; Stanford University(约翰斯·霍普金斯大学; 新加坡国立大学; 圣路易斯华盛顿大学; 伊利诺伊大学厄巴纳-香槟分校; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Red-TTT方法,在攻击过程中实时更新攻击者参数,利用策略梯度整合信息,显著提升越狱攻击成功率,从55.9%提升至72.4%。

AI 中文摘要

大型语言模型仍然容易受到越狱攻击,而自动红队测试是在大规模范围内发现大型语言模型中越狱漏洞的标准方法。当前的方法要么在测试时通过搜索、重写和树扩展来抽取更多样本,要么使用强化学习离线训练一个更强的攻击者。这两种方法都有一个共同的局限性:一旦针对特定目标行为的攻击开始,攻击者的权重就被冻结。它收集到的关于该行为的任何信号都保留在其上下文窗口中,并在之后被丢弃。攻击者从未在攻击过程中调整其提议分布,因此成功几乎完全取决于采样预算,而在大规模可负担的预算下,许多行为仍然无法被攻破。我们提出了Red-TTT,它在攻击每个行为的过程中更新攻击者的参数。在每一轮中,攻击者抽取一组候选样本,根据受害者的回复对其进行评分,并在抽取下一组之前执行一次策略梯度更新,从而将其发现的关于当前受害者的信息整合到权重中,而不是累积为上下文。我们还调整了训练目标以适应红队测试,其中成功是由单个最佳样本而非平均值来评判的。Red-TTT仅需要对受害者的采样访问权限,并且无需其他更改即可集成到现有的攻击流程中。与Best-of-N基线相比,Red-TTT在120个样本的预算下,平均攻击成功率从55.9%提高到72.4%,在每种配置下都优于基线,并破解了先前方法无法破解的许多行为。代码可在以下网址获取:https://this URL

英文摘要

Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with reinforcement learning. Both share a limitation: once an attack on a specific target behavior begins, the attacker's weights are frozen. Any signal it gathers about the behavior stays in its context window and is discarded afterward. The attacker never adapts its proposal distribution mid-attack, so success depends almost entirely on the sampling budget, and under a budget affordable at scale, many behaviors remain unbroken. We propose Red-TTT, which updates the attacker's parameters during the attack on each behavior. At each round, the attacker samples a group of candidates, scores them against the victim's replies, and takes a policy-gradient step before drawing the next group, so what it discovers about the current victim is consolidated into weights rather than accumulated as context. We also adapt the training objective to red teaming, where success is judged by the single best sample rather than the average. Red-TTT requires only sampling access to the victim and integrates into existing attack pipelines with no other changes. Against the Best-of-N baseline, Red-TTT raises attack success rate from 55.9\% to 72.4\% on average at a budget of 120 samples, improving over the baseline in every configuration and cracking many behaviors previous method cannot. The code is available at https://github.com/SaFo-Lab/Red-TTT

Comments29 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑