arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34415cs.LG

PROACT-Agent:用于实时安全的渐进式运行时监督与主动断路

PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety

Ding Jia, Wei Liu, Xianglong Du, Yingjie Li, Yingqing Yang, Huili Yu, Zhangsong Zhan, Chu Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

提出PROACT-Agent框架,通过渐进式轨迹展开、因果修正和文化感知数据本地化合成高保真轨迹,训练实时安全护栏,在双语基准上实现高精度检测并显著降低智能体攻击成功率。

中文摘要 AI 辅助

从大型语言模型(LLMs)向智能体的转变,将安全风险从有害文本提升至不可逆转的环境损害。虽然当前的防御措施大多仍是事后性的,但主动的运行时干预却因缺乏大规模、因果一致的数据而受阻。我们提出了PROACT-Agent,一个用于合成高保真轨迹以实现实时护栏的框架。我们识别出先前基准中一个关键的“安全漂移”问题,即宽松的标注范式未能强制执行时间一致性。PROACT-Agent通过以下方式解决这一问题:(1)渐进式轨迹展开,以揭示长上下文交互中隐藏的风险;(2)推理增强的因果修正,以强制单调因果一致性;(3)文化感知的数据本地化,以实现跨边界的鲁棒性。我们引入了PROACT-Bench,一个包含155,780个状态、通过多模型裁决标注的双语安全基准。在下次LLM推理之前评估更新后的上下文,所训练的护栏在完全源数据保留的情况下,实现了91.46%的不安全类别F1分数和90.63%的精确边界检测。在AgentDojo中,它将非拒绝服务型定向攻击成功率从20.82%降至0.40%。

英文摘要

The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PROACT-Agent, a framework for synthesizing high-fidelity trajectories to enable real-time guardrails. We identify a critical "safety drift" in prior benchmarks, where lenient annotation paradigms fail to enforce temporal consistency. PROACT-Agent addresses this through: (1) Progressive Trajectory Unrolling to reveal risks hidden in long-context interactions; (2) Reasoning-Augmented Causal Rectification to enforce monotonic causal consistency; and (3) Culturally-Aware Data Localization for cross-border robustness. We introduce PROACT-Bench, a bilingual safety benchmark with 155,780 states labeled through multi-model adjudication. Evaluating updated context before the next LLM inference, the trained guard achieves 91.46% unsafe-class F1 and 90.63% exact-boundary detection under complete source holdout. In AgentDojo, it reduces non-DoS targeted attack success from 20.82% to 0.40%.

发表机构

  • State Key Laboratory of Intelligent Vehicle Safety Technology, Changan Automobile(长安汽车智能车辆安全技术国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑