arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于仿真的LLM赋能的自主智能体验证中的多面向运行时验证

Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agents

Nikolaos Kekatos, Dimitrios Nikou, Anastasios Temperekidis, Alexios Lekidis, Nikolaos Kolokotronis, Panagiotis Katsaros, Stylianos Basagiannis

arXiv 2610.08928首次发表:更新:

发表机构

Clone Systems; International Hellenic University; University of Thessaly; University of Peloponnese; Aristotle University of Thessaloniki(Clone Systems公司; 国际希腊大学; 色萨利大学; 伯罗奔尼撒大学; 塞萨洛尼基亚里士多德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM智能体在国防决策中的合规性,提出多面向运行时验证框架,分解策略为空间/时间/语义三元组并融合判定,在预防性阻断下实现零攻击成功且无误报,有效防护多步违规。

AI 中文摘要

基于大语言模型(LLM)的智能体正进入国防参谋工作中的决策支持角色,在这些角色中,它们必须遵守的义务已经以书面形式确定并具有约束力,而且由于模型作为采购组件到达,重训练不可用作控制手段。可以置于工程控制之下的是智能体与其所作用的系统之间的接口。这些义务同时具有空间性、时间性和文本语义性,而违规通常存在于多步交互的组合中,这就是为什么逐事件防护栏会遗漏顺序工具攻击链的原因。我们提出一个多面向运行时验证框架,该框架将自然语言策略条款分解为一个规范事件流上的类型化空间/时间/语义三元组,使用各自的监控规范检查每个面向,并通过一个携带来源的四值代数融合判定结果。空间面向在一个加权双类型位置图上进行解释,在该图中,任务几何和信息发布拓扑是同一个对象;我们证明这些空间义务通常不被一阶时间规范所包含。过去时间面向在未经修改的MonPoly引擎上运行,该引擎在每个时间点都与我们的参考监控器一致。在两个任务领域(伤员后送和有争议的持续保障)以及一个民用领域中,在预防性阻断策略下的组合将攻击成功率降至零,没有观察到误报,且每事件成本为微秒级,而每个单独的面向和每一对面向都留下了相当大份额的攻击成功。在闭环实验中,一个对策略无知的规划器在大多数无防护任务中达到违规状态,而在有防护时则没有,并且五次拒绝事件中有四次仍能恢复到合规结果。

英文摘要

LLM-based agents are entering decision-support roles in defence staff work, where the obligations they must respect are already written down and binding, and where retraining is not available as a control because models arrive as procured components. What can be placed under engineering control is the interface between the agent and the systems it acts on. Those obligations are at once spatial, temporal and text-semantic, and a violation typically lives in the composition of a multi-step interaction, which is why per-event guardrails miss sequential tool-attack chains. We present a multi-aspect runtime-verification framework that decomposes a natural-language policy clause into a typed spatial/temporal/semantic triple over one canonical event stream, checks each aspect with its own monitoring specification, and fuses the verdicts through a four-valued algebra that carries provenance. The spatial aspect is interpreted over a weighted two-sorted location graph in which mission geometry and information-release topology are one object; we show that these spatial obligations are not in general subsumed by a first-order temporal specification. The past-time aspect runs on the unmodified MonPoly engine, which agrees with our reference monitor at every time point. Across two mission domains, casualty evacuation and contested sustainment, and one civil domain, composition under the precautionary blocking policy drives attack success to zero with no observed false positives and microsecond-scale per-event cost, while every single aspect and every pair leaves a substantial share of attacks succeeding. In a closed-loop experiment a policy-naive planner reaches a violating state in most unshielded missions and in none when shielded, and four refused episodes in five still recover to a compliant outcome.

Comments28 pages, 2 figures, 2 tables. Accepted at Modelling and Simulation for Autonomous Systems (MESAS 2026), to appear in Springer LNCS

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑