arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NEXUS:工具使用型大语言模型智能体的结构化运行时安全

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

Elias Hossain, Md Mehedi Hasan Nipu, Tasfia Nuzhat Ornee, Rajib Rana, Niloofar Yousefi

arXiv 2607.19356首次发表:更新:

发表机构

University of Central Florida; North South University; University of Southern Queensland(中佛罗里达大学; 南北大学; 南昆士兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究工具使用型大语言模型智能体运行时安全问题,提出NEXUS结构化计划安全监控器,结合多种方法进行分级升级。在多个基准测试中表现出色,如合成基准测试F1分数达0.949等,还具有低延迟、低开销特点,相关成果已公开。

AI 中文摘要

使用工具的大语言模型智能体越来越多地执行高影响操作,因此运行时安全监控至关重要。我们提出了NEXUS(神经执行实用程序与安全),这是一种结构化计划安全监控器,它应用正式的干预策略在四种操作中进行选择:允许、阻止、请求确认或请求修订。NEXUS结合了确定性安全规则、参数级检查和校准的逻辑回归风险评分以进行分级升级。在一个128实例的合成基准测试中,NEXUS的F1分数达到0.949,四类干预准确率为0.6406,比仅基于规则的干预选择高出27.3个百分点。在R-Judge上也优于仅基于规则的方法(F1 = 0.861对0.849),在AgentHarm上由于威胁模型限制与仅基于规则的方法相当,在IPI上99%控制允许时ASR为0%。在无规则的NEXUS-Stress基准测试中,NEXUS的F1分数达到0.881。NEXUS的中位延迟为0.205毫秒,给典型智能体循环增加的开销不到0.1%。代码、基准测试和校准风险评分器已公开发布。

英文摘要

Tool-using LLM agents increasingly execute high-impact actions, making runtime safety monitoring essential. We present NEXUS (Neural EXecution Utility and Safety), a structured-plan safety monitor that applies a formal intervention policy to select among four actions: allow, block, request confirmation, or request revision. NEXUS combines deterministic safety rules, argument-level inspection, and a calibrated logistic-regression risk score for graded escalation. On a 128-instance synthetic benchmark, NEXUS achieves an F1 score of 0.949 and a 4-class intervention accuracy of 0.6406, outperforming rule-only intervention selection by 27.3 percentage points. It also improves over rule-only on R-Judge (F1 = 0.861 vs. 0.849), matches rule-only on AgentHarm due to threat-model limits, and achieves 0% ASR at 99% control allow on IPI. On the rule-blind NEXUS-Stress benchmark, NEXUS reaches an F1 score of 0.881, highlighting the difficulty of fine-grained intervention routing. With 0.205 ms median latency, NEXUS adds under 0.1% overhead to typical agent loops. Code, benchmarks, and the calibrated risk scorer are publicly released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑