arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15994cs.CLcs.LG

潜在暗流:普通拼写错误如何破坏探针

Latent Undertow: How Ordinary Typos Break Probes

Elad David, Max Fomin, Amit LeVi

中文总结 AI 辅助

本研究揭示普通拼写错误会显著破坏基于隐藏状态的恶意提示探针,并提出KV缓存分叉方法,通过附加固定后缀利用扰动快速衰减,弥补95%的检测性能差距。

中文摘要 AI 辅助

大型语言模型(LLMs)能流畅地处理普通的打字变体:一个拼写错误或缺失的标点符号不会实质性地改变用户意图和模型的响应。然而,通过读取模型隐藏状态来检测恶意提示的探针却讲述了不同的故事:同样的编辑在受扰动的token处将读出向量旋转了43至56度,并在约10个下游token内衰减至15%以下。每条消息堆叠约3个常见拼写错误,会使单位置提示注入探针的TPR@FPR=1%降低12.0个百分点,这一差距仅靠重新校准无法弥补。多位置聚合能治愈局部扰动(损失<=0.5),但只能减弱分布式扰动,即使基于注意力和最大值的聚合器仍会下降约3.8个百分点。对于单位置探针,我们引入了KV缓存分叉:在用户消息后附加一个短的固定后缀,使探针能读取扰动下游的几个token,利用其快速的空间衰减。这弥补了95%的差距(残余-0.6个百分点),比扰动增强训练(-3.7个百分点)好一个数量级。旋转和衰减的几何结构在Llama-3.1-8B、Qwen3-8B和Gemma-4-E4B上复现;探针评估在Llama-3.1-8B上进行。代码:此HTTPS URL

英文摘要

LLMs handle ordinary typing variation fluently: a typo or missing punctuation leaves both user intent and the model's response substantively unchanged. Yet probes that detect malicious prompts by reading the model's hidden states tell a different story: the same edit rotates the readout vector by 43--56 at the perturbed token, decaying below 15% within ~10 downstream tokens. Stacking ~3 common typos per message cuts a single-position prompt-injection probe's TPR@FPR$=1% by 12.0pp, a gap recalibration alone cannot close. Multi-position aggregation cures localized perturbations (<= 0.5 loss) but only attenuates distributed ones, where even attention- and max-based aggregators still drop ~3.8pp. For single-position probes, we introduce a KV-cache fork: a short fixed suffix appended after the user message lets the probe read a few tokens downstream of the perturbation, exploiting its rapid spatial decay. This closes 95% of the gap (-0.6pp residual) -- an order of magnitude better than perturbation-augmented training (-3.7pp). The rotation-and-decay geometry replicates on Llama-3.1-8B, Qwen3-8B, and Gemma-4-E4B; probe evaluation is on Llama-3.1-8B. Code: https://github.com/eladd-ai/latent-undertow

补充信息

↑