arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18515cs.AIcs.LG

超越常规合规:巧妙数据培养大型语言模型的安全警觉性

Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

Youjia Wang, Lin Xu, Yang Sun, Yuxiao Lu, Chengfang Fang, Jie Shi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出通过巧妙问题训练大型语言模型,增强其对误导性推理的警觉性,从而提升对越狱攻击的鲁棒性,并降低平均攻击成功率,补充常规安全对齐。

中文摘要 AI 辅助

安全对齐教导大型语言模型(LLMs)识别有害请求并拒绝危险指令。然而,当有害意图隐藏在看似良性的上下文中时,对齐模型可能会失效。因此,稳健的安全既需要安全边界的知识,也需要警觉性:即检测异常前提、误导性推理以及表面语义之下的潜在风险的能力。警觉性要求模型在行动前审视请求的潜在意图和假设。为了培养这种能力,我们引入了巧妙问题,这些问题不一定与安全相关,但包含误导性前提、非典型推理或微妙的矛盾。我们假设,学会看穿此类推理陷阱可以迁移到安全关键场景。实验表明,巧妙训练提高了对分布外越狱攻击的鲁棒性,并增强了后续的安全微调。此外,将巧妙增强到现有的最先进安全对齐流程中,在我们评估的设置中确立了新的最先进水平,将九个骨干-基准组合的平均攻击成功率(ASR)从17.40%降低到15.05%。匹配安全微调后的痕迹分析表明,安全判断更有可能在有害规划开始之前主导响应。条件性理论分析进一步刻画了从巧妙数据中学到的不变性何时可以迁移到与安全相关的输入。这些发现表明,巧妙数据可以增强模型警觉性,并补充常规安全对齐。

英文摘要

Safety alignment teaches large language models (LLMs) to recognize harmful requests and reject risky instructions. Yet aligned models can fail when harmful intent is concealed within seemingly benign contexts. Robust safety therefore requires both knowledge of safety boundaries and \textbf{vigilance}: the ability to detect unusual premises, misleading reasoning, and latent risks beneath surface-level semantics. Vigilance requires models to scrutinize a request's underlying intent and assumptions before acting. To cultivate this capability, we introduce \textbf{cunning questions}, which are not necessarily safety-related but contain misleading premises, atypical reasoning, or subtle inconsistencies. We hypothesize that learning to look beyond such reasoning traps can transfer to safety-critical scenarios. Experiments show that Cunning training improves robustness to out-of-distribution jailbreak attacks and strengthens subsequent safety fine-tuning. Furthermore, augmenting an existing state-of-the-art safety alignment pipeline with Cunning establishes a new state of the art across our evaluated settings, reducing mean ASR across nine backbone--benchmark combinations from 17.40\% to 15.05\%. Trace analysis after matched safety fine-tuning suggests that safety judgments are more likely to govern responses before harmful planning begins. A conditional theoretical analysis further characterizes when invariance learned from cunning data can transfer to safety-related inputs. These findings suggest that cunning data can strengthen model vigilance and complement conventional safety alignment.

发表机构

  • Huawei(华为)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑