AI 中文总结
研究长期智能体安全潜在风险,提出JANUS框架,通过多智能体模拟合成轨迹,经预测与裁决耦合任务学习共享策略,用CoAA-RL联合优化,所产生的Vanguard模型提升了智能体安全防护及任务完成率。
AI 中文摘要
智能体安全正从内容审核转向在使用工具的智能体行动前预防操作失败。我们提出了JANUS,一个用于长期智能体安全的面向预见的框架,它训练防护器从部分轨迹预测延迟风险。JANUS通过多智能体模拟合成多样的智能体轨迹,并通过两个耦合任务学习共享策略:一个预测与安全相关未来的预测任务和一个根据观察到的前缀和预测未来判定安全的裁决任务。这两个任务用CoAA-RL联合优化,该方法根据预测对下游安全判断的效用进行奖励。由此产生的防护器模型Vanguard在执行前阻止不安全行动。在四个智能体安全基准测试中,Vanguard比基线防护器平均保护提高了15.9个百分点,同时良性任务完成率提高了5.1个百分点。
英文摘要
Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate delayed risks from partial trajectories. Janus synthesizes diverse agent trajectories via multi-agent simulation and learns a shared policy with two coupled tasks: an anticipation task that forecasts safety-relevant futures and an adjudication task that decides safety from both the observed prefix and anticipated future. The two tasks are jointly optimized with CoAA-RL, which rewards forecasts by their utility for downstream safety judgment. The resulting guard model, Vanguard, blocks unsafe actions before execution. Across four agent-safety benchmarks, Vanguard improves average protection by 15.9 percentage points over baseline guards while increasing benign task completion by 5.1 percentage points.