发表机构
Ryonix Labs Inc.(Ryonix实验室公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究揭示基于LTL的LLM智能体安全监视器覆盖差异源于攻击分布熵,提出熵-覆盖边界并验证,引入部署前熵测试,可预测监视器覆盖并支持架构感知选择。
AI 中文摘要
基于线性时序逻辑(LTL)和有限自动机(FSA)的运行时安全监视器正越来越多地被用于拦截LLM智能体中的不安全工具调用序列。然而,同一监视器在某些模型架构上实现了68%-75%的攻击覆盖率,而在其他架构上几乎为零,能力得分、训练数据或提示设计均无法对此作出解释。我们提供了缺失的理论:我们证明,任何固定不变量FSA监视器的召回率都存在上限,该上限由攻击分布的集中度决定,即k个最频繁触发-完成模式所覆盖的攻击比例。当攻击集中(香农熵低)时,小型固定不变量集可实现高召回率;当攻击分散在众多结构不同的模式中(熵高)时,无论不变量如何推导,任何可处理大小的固定不变量集都无法实现高召回率。我们在8种前沿LLM架构上验证了这一熵-覆盖边界:GPT类和DeepSeek后端产生高度集中的攻击(H≈0.24比特,1种模式覆盖96%),对应68%-75%的召回率;Gemini变体产生高熵分布(H≈2.81比特,7个聚类每个≤7%),对应近零召回率(6%-13%),且与架构匹配的再训练无关。熵解释了76%的覆盖方差(皮尔逊相关系数r=-0.87,p=0.005,95%置信区间[-0.98,-0.78]),在留一法下也成立(r在[-0.91,-0.82]之间)。我们引入了一种部署前熵测试,可通过小型攻击样本预测监视器覆盖情况,支持部署前的架构感知监视器选择。该边界和测试与架构无关,适用于任何基于FSA的离散动作序列运行时监视器。
英文摘要
Runtime safety monitors based on Linear Temporal Logic (LTL) and finite automata (FSA) are increasingly deployed to intercept unsafe tool-call sequences in LLM agents. Yet the same monitor achieves 68-75% attack coverage on some model architectures and near-zero on others, with no explanation from capability scores, training data, or prompt design. We provide the missing theory. We prove that the recall of any fixed-invariant FSA monitor is bounded above by the concentration of the attack distribution: the fraction of attacks covered by the k most frequent trigger-completion patterns. When attacks concentrate (low Shannon entropy), a small fixed invariant set achieves high recall; when they disperse across many structurally distinct patterns (high entropy), no fixed invariant set of tractable size can, regardless of how the invariants were derived. We validate this entropy-coverage bound across eight frontier LLM architectures. GPT-class and DeepSeek backends yield highly concentrated attacks (H ~ 0.24 bits; one pattern covers 96%), explaining 68-75% recall; Gemini variants yield high-entropy distributions (H ~ 2.81 bits; 7 clusters each <= 7%), explaining near-zero recall (6-13%), invariant to architecture-matched retraining. Entropy accounts for 76% of variance in coverage (Pearson r = -0.87, p = 0.005, 95% CI [-0.98, -0.78]), holding under leave-one-out (r in [-0.91, -0.82]). We introduce a pre-deployment entropy test that predicts monitor coverage from a small attack sample, enabling architecture-aware monitor selection before deployment. The bound and test are architecture-agnostic and apply to any FSA-based runtime monitor over discrete action sequences.
Comments6 pages, 1 figure. Accepted at the 13th IEEE International Conference on Intelligent Systems (IS'26), Varna, Bulgaria, 2026