arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

鲁棒评论家:防御大语言模型免受多轮攻击

Robust Critics: Defending LLMs Against Multi-Turn Attacks

Roman Belaire, Arunesh Sinha, Pradeep Varakantham

arXiv 2607.20472首次发表:更新:

AI 中文总结

研究针对大语言模型在多轮攻击下的安全问题,提出对话评论家引导采样(DCGS)框架,通过推断用户意图、建模为马尔可夫决策过程并学习价值和遗憾评论家来生成回复,在多个数据集上评估优于基线和前沿模型,还能迁移提升模型鲁棒性。

AI 中文摘要

当用户向语言模型询问有害内容时,这是真正的攻击还是误解但善意的问题?这种模糊性是大语言模型安全的核心挑战之一。假设最坏情况的模型会伤害合法用户,假设最好情况的模型则容易被利用。在多轮对话中问题更复杂,攻击者的真实意图可能在多次交流中逐渐显现,而现有安全框架采用上下文博弈处理,忽略对话轨迹。为此,我们提出对话评论家引导采样(DCGS)框架,通过在对话的每一轮推断用户意图来解决此问题。DCGS不是应用关于什么安全或不安全的固定规则,而是根据完整对话历史学习用户可能的意图并相应生成回复。我们将对抗性对话建模为马尔可夫决策过程,在单个令牌和话语(完整回复)级别学习基于价值和遗憾的评论家,通过动作价值评论家对候选回复进行评分。我们证明这种推理时的重新加权近似于基础策略的指数倾斜,保证在任何有限候选池中预期回报的提高,这是群体相关目标所不具备的属性。在CARES-18k、WildJailbreak、Redbench和Harmbench上进行评估,DCGS在对抗性对话任务上优于强大的鲁棒基线和前沿模型。DCGS还可迁移到前沿模型,无需微调即可提高其鲁棒性。

英文摘要

When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the central challenges of LLM safety. A model that assumes the worst harms legitimate users; one that assumes the best is easily exploited. The problem is compounded in multi-turn dialogue, where an attacker's true intent may only reveal itself gradually across many exchanges, yet existing safety frameworks apply a contextual bandit treatment, ignoring the trajectory of the conversation. To that end, we propose Dialogue Critic Guided Sampling (DCGS), a framework that addresses this by inferring user intent at every turn of dialogue. Instead of applying a fixed rule about what is or is not safe, DCGS learns what the user's intent is likely to be based on the full conversational history and generates responses accordingly. Formally, we model adversarial dialogue as a Markov Decision Process and learn value and regret-based critics at both the individual token and utterance (full response) levels, scoring candidate responses via an action-value critic. We prove that this inference-time reweighting approximates exponential tilting of the base policy, guaranteeing improvement in expected return for any finite candidate pool, a property that group-relative objectives do not exhibit. Evaluated on CARES-18k, WildJailbreak, Redbench, and Harmbench, DCGS outperforms strong robust baselines and frontier models on adversarial dialogue tasks. DCGS also transfers to frontier models, improving their robustness without fine-tuning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑