发表机构
National University of Singapore; MBZUAI; A*STAR Institute of Advanced Intelligence and Computing(新加坡国立大学; 穆罕默德·本·扎耶德人工智能大学; 新加坡科技研究局高级智能与计算研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示开放权重大语言模型中触发器标签滥用检测机制在对抗攻击下的脆弱性,提出统一攻击框架Untag,证明现有机制在输出转换或权重修改下完全失效。
AI 中文摘要
开放权重语言模型可以在开发者控制之外被下载、修改和部署,这限制了集中式安全防护措施的有效性。因此,近期研究提出了触发器标签机制,当模型在目标条件下被使用(例如生成钓鱼内容)时,该机制会产生可检测的信号。尽管这些机制借鉴了已有技术,但其在开放权重大语言模型中用于条件性滥用检测的应用相对较新。因此,现有研究工作尚未系统地研究触发器标签机制在对抗性攻击下的鲁棒性。为弥补这一空白,(i)我们形式化了触发器标签,并区分了令牌级触发器标签(在解码过程中引入水印启发信号)与权重级触发器标签(学习目标条件与可检测模型行为之间的后门启发关联)。此外,(ii)我们引入了Untag,一个统一的攻击框架,将机制特定的攻击面组织成通用分类体系。我们以钓鱼为案例研究,评估了代表性的令牌级和权重级触发器标签。我们发现,尽管触发器标签在受控环境中可能提供有用证据,但我们的攻击使现有触发器标签机制完全失效。因此,我们认为,当攻击者能够转换输出或修改开放权重时,这些机制不应被视为鲁棒的滥用检测器。
英文摘要
Open-weight language models can be downloaded, modified, and deployed beyond their developers' control, limiting the effectiveness of centrally enforced safeguards. Recent work has therefore proposed \emph{trigger-tag} mechanisms that produce a detectable signal when a model is used under a target condition, such as generating phishing contents. Although these mechanisms borrow from established techniques, their use for conditional misuse detection in open-weight LLMs is relatively new. Therefore, existing research works have not systematically studied the robustness of trigger-tag mechanisms under adversarial attacks. To close this gap, (i)~we formalize trigger-tags and distinguish \emph{token-level trigger-tags}, which introduce watermark-inspired signals during decoding, from \emph{weight-level trigger-tags}, which learn backdoor-inspired associations between target conditions and detectable model behavior. Furthermore, (ii)~we introduce \Untag, a unified attack framework that organizes their mechanism-specific attack surfaces into a common taxonomy. We evaluate representative token-level and weight-level trigger-tags using phishing as a case study. We find that while trigger-tags may provide useful evidence in controlled settings, our attacks render the existing trigger-tag mechanisms to be entirely ineffective. Consequently, we argue that these mechanisms should not be treated as robust misuse detectors when attackers can transform outputs or modify open weights.