DreamGuard:基于风险感知世界模型的LLM智能体高效运行时护栏
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
AI总结:
针对LLM智能体运行时护栏的长视野风险盲点问题,提出基于风险感知世界模型的主动式护栏DreamGuard,经实验验证其安全-效用权衡更优且延迟低。
AI中文摘要:
随着大语言模型(LLM)智能体越来越多地调用外部工具并与现实世界系统交互,不安全的操作可能会对外部状态、用户数据和下游服务造成不可逆的后果。近期的运行时护栏通过在执行前检查拟议操作来缓解此类风险,但许多护栏仍属于反应式:它们主要评估当前操作的表面安全性,缺乏对风险如何在轨迹中演变的显式模型。这一限制为长视野风险创造了关键盲点,即单独看似良性的操作可能会逐渐使智能体向危险状态漂移。针对这一问题,我们提出了DreamGuard,一种围绕风险感知世界模型构建的LLM智能体主动式护栏。该世界模型在轨迹上维护紧凑的循环隐状态,并预测未来隐状态,DreamGuard据此得出即时危险和前缀风险证据,随后将这些多视野信号融合为执行前的干预决策。在四个基准测试和一项在线护栏评估实验中,DreamGuard的表现优于通用、反应式和主动式护栏基线,在所有评估护栏中实现了最佳的安全-效用权衡,且每次调用的平均端到端延迟为25毫秒。
英文摘要:
As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services. Recent runtime guardrails mitigate such risks by checking proposed actions before execution, but many remain reactive: they primarily assess the apparent safety of the current action, lacking an explicit model of how risk evolves across the trajectory. This limitation creates a critical blind spot for long-horizon risks, where individually benign-looking actions can gradually drift the agent toward hazardous states. In response, we propose DreamGuard, a proactive guardrail for LLM agents built around a risk-aware world model. The world model maintains a compact recurrent latent state over the trajectory and predicts future latent states from which DreamGuard derives immediate-hazard and prefix-risk evidence. It then fuses these multi-horizon signals into intervention decisions before execution. Experiments across four benchmarks and an online guardrail evaluation show that DreamGuard outperforms generic, reactive, and proactive guardrail baselines, achieves the best safety-utility trade-off among evaluated guardrails, and maintains an average end-to-end latency of 25 ms per call.