arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分层大语言模型防御作为集成体:访问层级、推理成本与防御层间实测失败相关性

Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers

Abrar Alotaibi, Muhammad Shahid Jabbar, Sadam Al-Azani, Moataz Ahmed

arXiv 2608.28327首次发表:更新:

AI 中文总结

该研究针对LLM防御堆叠的假设,实测发现其防御层间存在正失败相关性,导致防御效果未达预期,需端到端实测堆叠表现而非仅依赖多样性选择成员。

AI 中文摘要

从业者通过堆叠防御措施来保护大语言模型(LLM),假设这些防御层会产生叠加效果。防御堆叠本质上是一个集成体,而集成体产生叠加效果的条件是LLM安全文献中推荐但从未被实测验证的:即各个防御层必须在不同输入上失效。两种工具使这一条件可被量化。对抗者访问层级模型(AATM)根据对抗者拥有的访问权限对其进行分级,范围从仅系统访问(A0)到对训练数据的影响(A4)。成本模型将防御措施分为五类推理时开销;由于其中两类需要训练权重或读取激活值,它们将防御者分为与AATM对抗者类似的层级。由此我们推导防御堆叠的表现,且防御者关心的指标出现分化:覆盖度在一个层级内达到饱和,成本随类别上升,虚假拒绝作为并集累积,仅在独立性条件下残余攻击成功率呈乘积下降。我们对独立性进行了实测:让一个自适应对抗者对抗七层防御堆叠,在15个可测量的防御层对中,失败相关性均为正(φ值从0.30到0.75),联合残余攻击成功率比乘积预测值高出最多0.172。按行为难度分层后,大部分关联消失,因此这种依赖性主要源于共同原因,但它在置换推理、多数投票 grader 标签和外部校准阈值下仍然存在。该防御堆叠对五分之四的良性提示进行了拒绝,同时与最强的单个防御层相比在统计上无差异。这种依赖性是架构性的而非采样性的:各层通过它们共同封装的模型产生关联,因此扩大成员池也无法削弱这种依赖性。因此,多样性可用于选择防御堆叠成员,但无法预测组装后的防御堆叠表现,必须端到端进行实测。

英文摘要

Practitioners defend large language models (LLMs) by stacking defenses, assuming the layers compound. A stack is an ensemble, and ensembles compound only under a condition the LLM security literature recommends but never measures: the members must fail on different inputs. Two instruments make that measurable. The Adversary Access-Tier Model (AATM) grades an adversary by the access it holds, from system-only (A0) to influence over training data (A4). A cost model sorts defenses into five classes of inference-time overhead; because two classes require training weights or reading activations, they tier the defender as AATM tiers the adversary. From these we derive how a stack behaves, and the quantities a defender cares about diverge: coverage saturates within a tier, cost rises by class, false refusals accumulate as a union, and residual attack success falls multiplicatively only under independence. We measure that independence. Running one adaptive adversary against a seven-layer stack, failure correlation is positive in all fifteen measurable pairs ($ϕ$ from $0.30$ to $0.75$), and the joint residual exceeds the multiplicative prediction by up to $0.172$. Stratifying on behavior difficulty dissolves most of the association, so the dependence is predominantly common-cause, but it survives permutation inference, majority-vote grader labels, and externally calibrated thresholds. The same stack refuses four in five benign prompts while remaining statistically indistinguishable from its strongest single layer. The dependence is architectural rather than sampling-based: members correlate through the model they all wrap, so no wider member pool weakens it. Diversity therefore selects stack members but does not predict what an assembled stack delivers, which has to be measured end to end.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑