发表机构
National University of Singapore; A*STAR Institute of Advanced Intelligence and Computing(新加坡国立大学; A*STAR高级智能与计算研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发现分词可被利用为侧信道,提出无参考攻击Toketive,绕过知识编辑与遗忘,检测并重建被抑制信息,在多个模型和数据集上显著优于现有基线。
AI 中文摘要
开放权重的大型语言模型(LLM)赋予下游用户对推理栈的控制权,但这种灵活性可能破坏发布后关于敏感知识已被修改或移除的保证。模型编辑和机器遗忘技术用于在不从头重新训练模型的情况下修改或移除目标知识。然而,现有针对这些技术的安全评估存在两个关键局限。首先,它们通常需要访问原始编辑前/遗忘前模型或辅助分类器,以检测修改或重建编辑前行为。其次,它们在输入的规范分词下评估修改,隐含地将分词视为良性预处理步骤。我们证明这一假设造成了安全漏洞:同一输入字符串可由多种有效的替代分词表示,这些分词会引发不同的计算轨迹,使攻击者能够绕过局部修改并恢复本应被抑制的信息。我们提出Toketive,一种简单而强大的无参考攻击,利用基于分词的侧信道来(i)检测被修改的知识并(ii)重建相应的编辑前响应。它仅作用于已发布的模型,既不需要编辑前模型、训练数据、影子模型,也不需要辅助分类器。在五个LLM、六个数据集以及六种编辑和遗忘技术中,我们发现38.6%的替代分词绕过了修改并恢复了编辑前响应。Toketive检测被修改事实的F1分数为84.2%,相对最强基线提升了26.2%,并以74.5%的top-5准确率重建编辑前响应,比最佳基线高出21.7%。我们的结果表明,未经对抗性评估的局部修改不应被视为稳健的知识控制边界,尤其是在替代表示存在的情况下。
英文摘要
Open-weight LLMs give downstream users control over the inference stack, but this flexibility can undermine post-release guarantees that sensitive knowledge has been modified or removed. Model editing and machine unlearning are used to modify or remove targeted knowledge without retraining models from scratch. However, existing security evaluations of these techniques face two critical limitations. First, they typically require access to either the original pre-edit/unlearning model or auxiliary classifiers to detect modifications or reconstruct pre-edit behavior. Second, they evaluate modifications under the canonical tokenization of an input, implicitly treating tokenization as a benign preprocessing step. We show that this assumption creates a security gap: the same input string can be represented by alternative valid tokenizations that induce different computational trajectories, allowing an adversary to bypass localized modifications and recover information intended to be suppressed. We introduce Toketive, a simple yet powerful reference-free attack that exploits the tokenization-based side channel to (i) detect modified knowledge and (ii) reconstruct the corresponding pre-edit response. It operates solely on the released model and requires neither the pre-edit model, training data, shadow models, nor auxiliary classifiers. Across five LLMs, six datasets, and six editing and unlearning techniques, we find that 38.6% of alternative tokenizations bypass the modification and recover the pre-edit response. Toketive detects modified facts with an F1 score of 84.2%, a 26.2% relative gain over the strongest baseline, and reconstructs pre-edit responses with 74.5% top-5 accuracy, 21.7% higher than the best baseline. Our results show that localized modifications should not be treated as robust knowledge-control boundaries without adversarial evaluation over alternative representations.