发表机构
Columbia University; University at Buffalo(哥伦比亚大学; 布法罗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出环境引导方法,通过将智能体执行状态建模为数据库表并运行时检查数据流策略,在违规时提供反馈引导安全轨迹,在AgentDyn上实现0%攻击成功率并提升任务成功率。
AI 中文摘要
LLM智能体即使被指示安全行事,也可能进行不安全的工具调用。现有防御措施在智能体执行前进行约束、修改工具输入/输出,或依赖LLM评判器;这些方法可能依赖于模型行为,或在不帮助智能体恢复的情况下阻止不安全操作。我们认为执行环境应在智能体运行时强制安全,并在违规发生时引导其转向安全替代方案——我们称之为环境引导。我们通过将智能体和执行环境状态建模为数据库表、跟踪记录级数据流,并在运行时对照声明性策略检查这些数据流来实现这一点。当检测到违规时,策略和上下文特定的反馈引导智能体走向安全轨迹。在AgentDyn上,这使得智能体在实现0%攻击成功率的同时,相比无防御提高了任务成功率。
英文摘要
LLM agents can make unsafe tool calls even when instructed to behave safely. Existing defenses constrain agents before execution, modify tool inputs/outputs, or rely on LLM judges; these approaches may depend on model behavior or block unsafe actions without helping the agent recover. We argue that the execution environment should instead enforce safety as the agent runs and steer it toward safe alternatives when violations occur---we call this Environment Steering. We implement this by modeling the agent and harness execution state as database tables, track the record-level data flows, and check these data flows against declarative policies during runtime. When violations are detected, policy- and context-specific feedback steers the agent toward safe trajectories. On AgentDyn, this enables the agent to improve task success rate over no-defense while achieving 0% attack success rate.
Comments9 pages, 11 figures, REALM Workshop, EMNLP 2026