arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

虚假预言:论智能体系统中世界模型的安全性

False Prophets: On the Security of World Models in Agentic Systems

Erik Imgrund, Anna Wimbauer, Klim Kireev, Konrad Rieck

arXiv 2607.23147首次发表:更新:

发表机构

BIFOLD; TU Berlin(BIFOLD; 柏林工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究智能体系统中世界模型的安全性,发现其存在特定漏洞,引入安全基准数据集,指出攻击者能诱导错误预测,成功率达95%,并为从业者提供减轻危害、强化系统的实用建议。

AI 中文摘要

大型语言模型如今为能够在不同环境中执行复杂多步任务的自主智能体提供支持。准确可靠地执行这些任务需要智能体预测其行动结果。近期研究提议通过经过特殊训练的环境模拟器——世界模型来增强预测能力。虽然世界模型能提升性能,但也可能误导智能体执行有害行动,带来重大安全和隐私风险。本文提出了关于在智能体系统中使用世界模型的安全担忧。我们发现了一系列特定于世界模型的漏洞,这些漏洞可在基于终端的智能体中被利用来执行恶意代码或提取敏感数据。为促进未来发展,我们引入了一个为基于文本的世界模型设计的安全基准数据集。我们认为一些风险是近似世界建模所固有的,并表明攻击者能在智能体管道中诱导错误预测,成功率高达95%,这可能导致意外命令执行、拒绝服务、钱包资金流失和私人信息提取。最后,我们为从业者提供了减轻已发现危害并强化智能体系统的实用建议。

英文摘要

Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable execution of these tasks requires the agent to predict the results of its actions. Recent research proposes to enhance predictive capabilities via specially trained environment simulators-world models. While world models can improve performance, they can also mislead agents into executing harmful actions, creating significant security and privacy risks. In this paper, we raise security concerns regarding the usage of world models in agentic systems. We discover a range of world model specific vulnerabilities, which can be exploited in terminal-based agents to execute malicious code or extract sensitive data. To facilitate future development, we introduce a security benchmark dataset designed for text-based world models. We argue that some risks are intrinsic to approximate world modeling, and show that attackers can induce mispredictions in agentic pipelines with up to 95% success rate, possibly resulting in unintended command execution, denial of service, drainage of wallet and private information extraction. Finally, we provide practical recommendations for practitioners to mitigate the discovered harms and harden agentic systems.

Journal ref19th Workshop on Artificial Intelligence and Security (AISEC), 2026

DOI:10.1145/3847352.3848097

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑