当世界说谎:针对潜在世界模型的下游控制后门攻击
When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control
- Radboud University(拉德堡德大学)
- IKERLAN Technology Research Centre(IKERLAN技术研究中心)
- University of Bergen(卑尔根大学)
- University of Zagreb(萨格勒布大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究揭示预训练世界模型重用中的供应链后门攻击,通过潜在状态路由与动力学重塑,使受害者优化自行发现攻击者目标动作,劫持下游控制器,且能通过干净数据诊断。
中文摘要 AI 辅助
预训练世界模型,即学习型模拟器,将观测编码为潜在状态并预测其在动作下的演化,正开始像预训练编码器和语言模型今天被重用一样,被用作现成的控制动力学骨干。我们表明,这种重用打开了一个供应链后门:仅控制一个发布检查点的对手可以劫持下游控制器,即使受害者在完全干净的数据上训练和评估,且从未看到触发器。该攻击不编码显式的触发器到动作规则。相反,被投毒的模型将携带触发器的观测路由到选定的潜在区域,并重塑那里的局部动力学,使得受害者自身的优化(Dreamer风格的想象中演员训练,或基于预测未来的MPC/CEM规划)自行重新发现攻击者的目标动作。在多个控制任务和触发器家族中,触发器将控制器的动作引向攻击者的目标,控制每个动作维度,并在最强设置下劫持100%的触发步骤。该检查点仍能通过受害者在部署前会运行的干净数据诊断,干净任务成功率至少保持约75%。该效果在时间上受门控:仅在触发器存在时出现,触发器移除时消失。对触发器盲的修复依赖于预算:适度的干净微调可以保持干净效用,同时使触发失败保持完整,而足够激进的适应只能在大幅降低干净控制后移除它。因此,世界模型骨干本身是控制的一个新兴且未充分研究的攻击面。完整代码和工件可在我们的仓库中获取。
英文摘要
Pretrained world models, learned simulators that encode an observation into a latent state and predict how it evolves under actions, are beginning to be reused as off-the-shelf dynamics backbones for control, like pretrained encoders and language models are reused today. We show that this reuse opens a supply-chain backdoor: an adversary who controls only a released checkpoint can hijack the downstream controller, even though the victim trains and evaluates entirely on clean data and never sees the trigger. The attack encodes no explicit trigger-to-action rule. Instead, the poisoned model routes trigger-bearing observations into a chosen latent region and reshapes the local dynamics there, so that the victim's own optimization (Dreamer-style actor training in imagination, or MPC/CEM planning over predicted futures) re-discovers the attacker's target action on its own. Across several control tasks and trigger families, the trigger steers the controller's action toward the attacker's target, controlling every action dimension and hijacking 100\% of triggered steps on the strongest settings. The checkpoint still passes the clean-data diagnostics a victim would run before deployment, with clean-task success retaining at least $\sim$75\%. The effect is temporally gated: it appears only while the trigger is present and disappears when the trigger is removed. Trigger-blind repair is budget-dependent: moderate clean fine-tuning can preserve clean utility while leaving the triggered failure intact, whereas sufficiently aggressive adaptation can remove it only after substantially degrading clean control. The world-model backbone itself is therefore an emerging and underexamined attack surface for control. The full code and artifacts are available in our repository.