发表机构
Clone Systems; International Hellenic University; University of Thessaly; Aristotle University of Thessaloniki(Clone Systems公司; 国际希腊大学; 色萨利大学; 塞萨洛尼基亚里士多德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估MFOTL在离线重放基准轨迹上的形式化监控,发现其能有效标记攻击但误报率高,且溯源上下文可被攻击,通过绑定溯源可消除规避,并提出了十二字段的轨迹模式。
AI 中文摘要
针对工具使用型LLM智能体的防护措施通常是特定于应用的规则,这使得多步骤、数据相关的安全策略难以指定、审计和重用。作为一种声明式替代方案,我们评估了度量一阶时序逻辑(MFOTL),通过未经修改的MonPoly监视器离线重放AgentDojo、STAC和R-Judge已记录的轨迹,无需运行智能体。在这些语料库上,五个通用义务标记了71.8%的STAC攻击链和70.1%的成功的AgentDojo攻击,但也在29.3%的良性运行上触发。这种不精确性源于语料库而非逻辑:它们很少记录批准,且从不记录时间戳,因此依赖历史的义务退化为检测风险行为类型。当轨迹确实携带关系上下文时,具有溯源感知的策略能更好地区分;然而,该上下文本身是可攻击的,一条植入的记录就能在94-99%的本会被标记的运行上击败朴素的溯源检查。将溯源绑定到产生它的查找操作上,可以在不损失检测率或良性触发率的情况下消除这种规避。综合来看,这些结果表明,当轨迹暴露可信历史时,形式化时序监控才真正有价值。因此,我们量化了当前基准与这一目标的差距,并提出一个十二字段的、可用于执行的轨迹模式。
英文摘要
Guardrails for tool-using LLM agents are usually application-specific rules, which makes multi-step, data-dependent safety policies hard to specify, audit and reuse. As a declarative alternative, we evaluate metric first-order temporal logic (MFOTL), replaying the recorded trajectories that AgentDojo, STAC and R-Judge already ship through the unmodified MonPoly monitor, offline and without running an agent. On these corpora, five generic obligations flag 71.8% of STAC attack chains and 70.1% of successful AgentDojo attacks, but also fire on 29.3% of benign runs. This imprecision stems from the corpora rather than the logic: they rarely record approvals and never record timestamps, so history-dependent obligations reduce to detecting risky action types. Where the trace does carry relational context, provenance-aware policies discriminate better; that context, however, is itself attackable, and one planted line defeats a naive provenance check on 94-99% of the runs it would otherwise flag. Binding provenance to the lookup that produced it closes this evasion at no cost in detection or benign firing. Taken together, these results show that formal temporal monitoring adds value exactly when the trace exposes trustworthy history. We therefore quantify how far current benchmarks are from that point and propose a twelve-field enforcement-ready trace schema.
Comments14 pages, 1 figure, 13 tables. Accepted at the 8th Workshop on CPS&IoT Security and Privacy (CPSIoTSec '26), co-located with ACM CCS 2026, The Hague