arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07051cs.LGcs.CR

TrojanWorld:通过想象引导对世界模型智能体进行后门攻击

TrojanWorld: Backdooring World-Model Agents via Imagination Steering

Wenkai Huang, Siyuan Liang, Gaolei Li, Yiming Li, Tianhao Peng, Jianhua Li, Dacheng Tao

首次发表
浏览论文内容

中文总结 AI 辅助

TrojanWorld通过物理触发器引导世界模型智能体的内部想象,实现隐蔽后门攻击,在保持高保真性能的同时诱导恶意行为。

中文摘要 AI 辅助

世界模型日益成为基于模型的强化学习智能体的预测核心,使智能体能够在行动前模拟未来动态并推理想象轨迹。其巨大的训练需求使得预训练世界模型在分发和复用方面具有吸引力,但也使下游系统面临模型供应链威胁。后门攻击提供了一种有针对性的隐蔽方式来利用此类供应链,然而其对交互式世界模型智能体的威胁在很大程度上仍未得到探索。为填补这一空白,我们提出了TrojanWorld,一个针对世界模型智能体的后门框架,通过引导内部想象来诱导攻击者指定的行为。场景中放置的物理对象作为触发器,使得在部署时可通过智能体自身的观测管道激活攻击,而无需对观测流进行数字篡改。为实现有效、隐蔽且持久的控制,TrojanWorld结合了决策反射归纳(利用决策反馈将触发器条件下的想象引导至攻击者指定的动作)、干净行为锚定(保持无触发器时的预测和行为保真度)以及因果传播(在触发器消失后沿后续轨迹维持诱导的偏好)。这些机制共同建立了一条从物理感知经腐败想象到恶意动作选择的端到端攻击链。在DeepMind Control、MetaWorld、MyoSuite和RoboDesk基准上使用TD-MPC2、DreamerV3和R2-Dreamer系统进行的实验表明,在触发器激活下,TrojanWorld实现了低至0.026的目标动作偏差,同时保留了至少98.8%的相应干净性能。即使在触发器移除后,被攻陷的智能体仍可能被困在诱导的行为轨迹中,继续执行攻击者指定的动作。

英文摘要

World models increasingly serve as the predictive core of model-based reinforcement learning agents, enabling them to simulate future dynamics and reason over imagined trajectories before acting. Their substantial training demands make pretrained world models attractive for distribution and reuse, exposing downstream systems to model supply chain threats. Backdoor attacks offer a targeted and stealthy means of exploiting such supply chains, yet their threat to interactive world-model agents remains largely unexplored. To fill this gap, we present TrojanWorld, a backdoor framework for world-model agents that induces attacker-specified behavior by steering internal imagination. A physical object placed in the scene acts as the trigger, enabling deployment-time activation through the agent's native observation pipeline without digitally manipulating the observation stream. To achieve effective, stealthy, and persistent control, TrojanWorld combines Decision-Reflective Induction to steer trigger-conditioned imagination toward attacker-specified actions using decision feedback, Clean Behavior Anchoring to preserve trigger-free predictive and behavioral fidelity, and Causal Propagation to sustain the induced preference along subsequent trajectories after the trigger disappears. Together, these mechanisms establish an end-to-end attack chain from physical perception through corrupted imagination to malicious action selection. Experiments with the TD-MPC2, DreamerV3, and R2-Dreamer systems across the DeepMind Control, MetaWorld, MyoSuite, and RoboDesk benchmarks show that under trigger activation, TrojanWorld achieves a target-action deviation as low as 0.026 while retaining at least 98.8% of the corresponding clean performance. Even after trigger removal, the compromised agent can remain trapped in the induced behavioral trajectory, continuing to execute attacker-specified actions.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑