发表机构
Institute of Automotive Technology, Technical University of Munich(慕尼黑工业大学汽车技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究基于大型视频扩散主干的驾驶世界模型难以控制的问题,提出通过可微能量函数在采样时操控规划轨迹,无需重新训练主干,以Open-Sora 2.0 MM-DiT主干模型为例验证,指出跨流耦合是端到端可控关键。
AI 中文摘要
基于大型视频扩散主干构建的驾驶世界模型可生成逼真场景,但难以控制:实施交通规范通常意味着重新训练主干或根据手工布局进行条件设定。我们探讨可控性是否真的需要训练。实验表明,联合生成未来视频和规划自我轨迹的整流流驾驶世界模型,可在采样时通过编码驾驶规范的可微能量函数完全操控规划轨迹,而无需对扩散主干进行特定知识的重新训练。具体而言,我们证明基于Open-Sora 2.0 MM-DiT主干构建的世界模型可通过在采样时注入能量引导,在反事实目标处制动。然而,我们发现生成的视频尚未通过主干的联合自注意力遵循操控轨迹,并确定跨流耦合是端到端可控推出的关键要求。
英文摘要
Driving world models built on large video-diffusion backbones generate realistic scenes but are hard to control: enforcing a traffic norm typically means retraining the backbone or conditioning it on hand-built layouts. We ask whether controllability requires training at all. Our experiment shows that a rectified-flow driving world model, which jointly generates future video and a planned ego trajectory, can have its planned trajectory steered entirely at sampling time by differentiable energy functions that encode driving norms, without knowledge-specific retraining of the diffusion backbone. Concretely, we demonstrate that a world model built on Open-Sora 2.0 MM-DiT backbone can be steered to brake at a counterfactual target by injecting energy guidance at sampling time. However, we find that the generated video does not yet follow the steered trajectory through the backbone's joint self-attention and identify the cross-stream coupling as a crucial requirement for end-to-end-controllable rollouts.
CommentsAccepted to Robotics: Science and Systems 2026 Robot World Models Workshop