发表机构
University of Technology Sydney, Australia(澳大利亚技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究语言模型逃避攻击在对抗性微调中的情况,利用检测器漏洞不对称性,引入新攻击家族,其愚弄率更高且能保持自然度,实验表明现有对策无效,揭示检测器在结构分布外转变下有持续漏洞并助力竞赛表现。
AI 中文摘要
我们研究了哪些语言模型逃避攻击能在最先进的对抗性微调中幸存下来,并制定了在2026年ELOQUENT Voight-Kampff排行榜上占据前5名的策略。虽然对抗性微调轻易地封堵了2025年获胜的逃避方法,但我们发现了检测器漏洞中的一个基本不对称性:将生成的文本推离检测器的训练分布能可靠地击败对抗性检测,而将其拉进分布(如模仿人类训练数据)则完全失败。利用这一点,我们引入了两个新的分布外攻击家族——跨十年寄存器攻击和现代主义意识流形式。这两种策略都能轻松绕过对抗性封堵,在保持自然度的同时,愚弄率比以前的方法高出约50倍。此外,实验表明,明显的部署者对策(用时代散文扩充训练数据)无法消除漏洞。我们的发现表明,包括经过对抗性微调的检测器家族,在结构分布外转变下存在持续漏洞,这一机制直接助力了我们在主要竞赛中的表现。
英文摘要
We investigate which language model evasion attacks survive state-of-the-art adversarial fine-tuning, developing strategies that sweep the top 5 positions on the ELOQUENT 2026 Voight-Kampff leaderboard. While adversarial fine-tuning trivially closes the 2025 winning evasion recipes, we uncover a fundamental asymmetry in detector vulnerability: pushing generated text out of the detector's training distribution reliably defeats adversarial detection, whereas pulling it into the distribution (e.g., mimicking human training data) fails completely. Exploiting this, we introduce two novel out-of-distribution attack families - cross-decade register attacks and modernist stream-of-consciousness form. Both strategies easily bypass adversarial closure, achieving up to approximately 50x higher fool rates than previous methods while preserving naturalness. Furthermore, experiments show that the obvious deployer countermeasure (augmenting training data with period prose) fails to close the vulnerability. Our findings show that the tested detector families, including adversarially fine-tuned ones, exhibit persistent vulnerabilities under structural out-of-distribution shifts, a mechanism that directly powers our leading competition performance.
CommentsCLEF2026, ELOQUENT2026