arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06612cs.CR

MechAudit-40:跨40种LLM攻击机制的白盒审计

MechAudit-40: White-Box Auditing across 40 LLM Attack Mechanisms

  • Saint Louis University(圣路易斯大学)
  • Washington University in St. Louis(华盛顿大学圣路易斯分校)

机构由 AI 辅助整理,请以论文原文为准。

Zhen Guo, Shanghao Shi, Shamim Yazdani, Ning Zhang, Reza Tourani

AI总结:

提出MechAudit-40测试平台和MechAudit审计器,利用内部表示偏移跨40种攻击机制实现白盒审计,在零预言机下高召回低误报,避免覆盖崩溃。

AI中文摘要:

虽然针对LLM的攻击涵盖了提示优化、多轮上下文操纵、检索投毒和模型后门,但白盒防御通常仅在孤立的攻击族上进行评估。因此,异质攻击是否会产生可泛化到未见威胁机制的内部表示偏移,仍然未知。我们提出了MechAudit-40,这是对五种开放权重模型架构上40种攻击机制的系统性评估。威胁特定成功标准、100,000个匹配的干净-攻击表示对、预定义类别和分组留出集,将真正的攻击诱导位移与目标规模、语料库偏差和数据泄漏捷径区分开来。在该测试平台上,攻击诱导出结构化的多深度轨迹,而非孤立的层尖峰。虽然原始峰值在不同架构间不可移植,但目标校准的轮廓保留了可迁移的几何签名:在完全机制留出下,仅凭隐藏状态就能以82.5%的准确率恢复未见攻击的威胁类别。受此发现指导,我们设计了MechAudit,一个运行时审计器,在严格零预言机约束下运行,无需干净基线轨迹或攻击元数据。MechAudit在0.70%的假阳性率下检测到81.1%的留出攻击执行,并在整个功能类别被 withheld 时保持78.1%的召回率。在匹配比较中,MechAudit是唯一避免机制级覆盖崩溃的检测器,在全部40种机制上保持超过50%的召回率。因此,内部表示支持针对校准良性参考的跨机制攻击暴露审计,但与下游任务受损和参数完整性脱钩。

英文摘要:

While LLM attacks span prompt optimization, multi-turn context manipulation, retrieval poisoning, and model backdoors, white-box defenses are typically evaluated on isolated attack families. Consequently, whether heterogeneous attacks leave internal representation shifts that generalize to unseen threat mechanisms remains unknown. We present MechAudit-40, a systematic evaluation of 40 attack mechanisms across five open-weight model architectures. Threat-specific success criteria, 100,000 matched clean-attack representation pairs, predefined categories, and grouped holdouts isolate genuine attack-induced displacement from target scale, corpus bias, and data-leakage shortcuts. Across this testbed, attacks induce structured multi-depth trajectories rather than isolated layer spikes. While raw peaks are non-portable across architectures, target-calibrated profiles preserve transferable geometric signatures: under complete mechanism holdout, hidden states alone recover the threat category of unseen attacks with 82.5% accuracy. Guided by this finding, we design MechAudit, a runtime auditor that operates under strict zero-oracle constraints without requiring clean baseline traces or attack metadata. MechAudit detects 81.1% of held-out attack executions at a 0.70% false-positive rate and maintains 78.1% recall when an entire functional category is withheld. In matched comparisons, MechAudit is the only detector that avoids mechanism-level coverage collapse, maintaining over 50% recall across all 40 mechanisms. Internal representations thus support cross-mechanism attack-exposure auditing against calibrated benign references, but decouple from downstream task compromise and parameter integrity.

补充信息

↑