arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37970cs.ROcs.CV

PhysWAM:用于自动驾驶的物理一致世界动作模型

PhysWAM: Physically Consistent World Action Model for Autonomous Driving

Dhruv Parikh, Fengcheng Yu, Quankai Gao, Jiawei Yang, Junjie Ye, Maulik Bhatt, Thang Vu, Charles Ochoa, Rowan McAllister, Igor Vasiljevic, Rajgopal Kannan, Vikt… 展开作者

Dhruv Parikh, Fengcheng Yu, Quankai Gao, Jiawei Yang, Junjie Ye, Maulik Bhatt, Thang Vu, Charles Ochoa, Rowan McAllister, Igor Vasiljevic, Rajgopal Kannan, Viktor Prasanna, Vitor Guizilini, Yue Wang

首次发表
浏览论文内容

中文总结 AI 辅助

PhysWAM提出统一世界动作模型,通过耦合点投影施加几何约束,联合生成视频、深度和自车运动,实现强规划性能与零样本迁移。

中文摘要 AI 辅助

世界动作模型(WAMs)联合预测场景将如何演变以及智能体应如何行动,然而,仅进行联合生成并不必然在这些预测上施加共享的几何约束。我们提出了PhysWAM,一个用于自动驾驶的统一世界动作模型,它在单个流匹配Transformer内共同去噪多视角视频、度量深度和自车运动。为了将世界和动作生成锚定在测量的场景几何上,我们引入了耦合点投影(CPP),它将生成的深度反投影为3D点,使用生成的SE(3)自车运动对这些点进行变换,并最小化它们与使用记录的自车运动变换的LiDAR点之间的距离。这一几何约束通过将生成的深度和运动与其标准流匹配目标一起进行联合监督,促进了与测量场景的物理一致性。在推理时,轨迹选择仅依赖于一个简单的无标签共识规则,无需学习评分器或模拟器反馈。我们在NAVSIM v1和v2规划、零样本闭环迁移以及未来视频和度量深度预测上评估了PhysWAM。尽管PhysWAM的选择过程简单,它仍实现了强大的规划性能,并零样本迁移到未见过的驾驶环境。它还生成了准确的度量深度和时间连贯的视频,其中CPP同时改善了规划和深度预测。这些结果共同表明,场景深度与自车运动之间的几何关系为在简单统一模型内耦合世界和动作生成提供了一种直接途径。

英文摘要

World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated $\mathrm{SE}(3)$ ego motion, and minimizes their distance to LiDAR points transformed using the recorded ego motion. This geometric constraint promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, trajectory selection relies only on a simple label-free consensus rule, with no learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM's simple selection procedure, it achieves strong planning performance and transfers zero-shot to unseen driving environments. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.

发表机构

  • University of Southern California(南加州大学)
  • Woven by Toyota(丰田编织公司)
  • Toyota Research Institute(丰田研究院)
  • DEVCOM Army Research Office(DEVCOM陆军研究办公室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑