发表机构
Delft University of Technology(代尔夫特理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对基于Gumbel的LLM推理验证防御,发现控制提示分布的攻击者可扩大容许token集合,使隐蔽信道泄露能力提升,建议动态校准抖动宽恕阈值。
AI 中文摘要
基于Gumbel的推理验证通过仅宽恕那些看似由诚实GPU非确定性产生的token选择,来限制大语言模型(LLM)的权重外泄,在良性提示流量下,隐写术攻击者的速度会降低200倍以上。该边界假设攻击者是被动的;我们证明,对于控制提示分布的攻击者,该边界会急剧恶化。由于验证器的容许token集合大小由模型自身的输出熵决定,因此,旨在破坏语法和子词结构的提示(而非良性对话流量)会扩大该集合,从而打开一个大得多的隐蔽信道。在6个参数规模从10亿到320亿的指令调优模型和3个随机种子上,我们最强的攻击(字符级和脚本级破坏)使每个token泄露的比特数相对于良性提示大致翻倍,将减速因子降至60倍至118倍。这些结果表明,针对该防御措施的静态、良性流量校准阈值是不够的,抖动宽恕阈值应改为针对局部token熵进行动态校准。
英文摘要
Gumbel-based inference verification bounds LLM weight exfiltration by only forgiving token choices that plausibly arise from honest GPU nondeterminism, reporting a >200x slowdown for a steganographic adversary under benign prompt traffic. This bound assumes a passive attacker; we show it degrades sharply against an adversary who instead controls the prompt distribution. Because the verifier's admissible-token-set size is driven by the model's own output entropy, prompts engineered to break grammatical and sub-word structure -- rather than benign conversational traffic -- widen that set and open a materially larger covert channel. Across six instruction-tuned models spanning 1B to 32B parameters and three random seeds, our strongest attack (character- and script-level disruption) roughly doubles bits leaked per token relative to benign prompts, cutting the slowdown factor to 60x - 118x. These results indicate that static, benign-traffic-calibrated thresholds are insufficient for this defense, and that jitter-forgiveness thresholds should instead be calibrated dynamically against local token entropy.
Comments4 pages, 1 figure, 1 table