LLM权重外泄的任意时刻有效检测
Anytime-valid detection of LLM weight exfiltration
查看机构详情
- MATS Research(MATS 研究)
- NASK – National Research Institute, MATS Research(NASK - 国家研究所,MATS 研究)
- Anthropic
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对LLM权重外泄问题,提出提示级e过程,结合跨响应弱证据,实现任意时刻误报控制,在四种模型上验证了对种子盲和种子感知攻击的检测效果,分析了信道容量与可检测性的权衡。
中文摘要 AI 辅助
被入侵的大语言模型(LLM)推理服务器可通过在看似合理的token选择中编码有效载荷位来泄露模型权重。在可信服务器中重放相同提示可暴露此类偏差,但良性数值非确定性也会导致token不匹配。因此,耐心的攻击者可隐藏在正常变异中,除非跨响应组合证据。我们引入一种提示级e过程,该过程基于可信良性流量校准全响应不匹配事件,并在顺序累积证据的同时,在校准转移假设下,控制无限监测时间内的任何误报概率。我们在四个模型上针对种子盲攻击和更强的种子感知攻击(该攻击仅在接近平局处隐藏有效载荷位以保持隐蔽)进行评估,分析信道容量与可检测性的权衡。与硬token级警报相比,e过程可跨响应组合弱证据,同时提供明确的任意时刻误报控制。
英文摘要
A compromised LLM inference server can leak model weights by encoding payload bits in otherwise plausible token choices. A replay of the same prompt in a trusted server can expose such deviations, but benign numerical nondeterminism also causes token mismatches. Patient attackers can therefore hide within normal variation unless evidence is combined across responses. We introduce a prompt-level e-process that calibrates whole-response mismatch events on trusted benign traffic and accumulates evidence sequentially while, under a calibration-transfer assumption, controlling the probability of any false alarm over an unbounded monitoring horizon. We evaluate it on four models against a seed-blind attack and a stronger seed-aware attack that hides payload bits only in near-ties to remain stealthy, analyzing the channel capacity vs detectability trade-off. Compared with a hard per-token alarm, the e-process combines weak evidence across responses while providing explicit anytime false-alarm control.