arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PV-WM:用于行人-车辆联合推演的异构微观-宏观世界模型

PV-WM: A Heterogeneous Micro-Macro World Model for Articulated Pedestrian-Vehicle Co-Rollout

Haozhuang Chi, Jingsong Liang, Ziying Song, Lei Yang, Shihao Li, Haoruo Zhang, Chen Lv

arXiv 2609.07328首次发表:更新:

发表机构

Nanyang Technological University; Beijing Institute of Technology(南洋理工大学; 北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PV-WM提出异构微观-宏观世界模型,循环推进行人关节运动与车辆状态,在Waymo基准上显著降低误差并减少计算开销。

AI 中文摘要

局部行人-车辆预测跨越异构物理尺度:行人结合根运动与关节运动,而车辆是由运动学状态和定向范围描述的刚体。现有的道路智能体预测器通常忽略行人关节运动,而姿态预测器则将车辆未来状态排除在学习推演之外。我们提出PV-WM,一个基于结构化感知后轨迹的仅历史世界模型。它在同步异构状态中循环推进行人根运动、15关节关节运动以及学习到的车辆状态。生成的行人块和车辆块提供下一个循环边界;车辆框由预测的中心和朝向结合观测范围重建,并且在每次转换后重新计算P-V几何关系。相对于匹配的单次完整状态预测器,循环执行将Root ADE降低12.7%,MPJPE降低14.8%。反馈干预表明,后续预测依赖于生成关节运动的内容、时间顺序和行人身份。在824个对齐的Waymo场景中,其中797个提供有效的未来车辆支持,相对于验证集选择的模块化专家,PV-WM将Root ADE降低5.2%,MPJPE降低7.6%,P-V距离误差降低11.9%,定向框最近接近误差降低5.8%。该单网络模型使用少57.1%的参数,每个局部场景平均FLOPs降低96.5%,实测p95延迟降低25.5%。PV-WM统一了这种异构未来状态,同时保留了类型特定的行人和车辆动力学。

英文摘要

Local pedestrian-vehicle forecasting spans heterogeneous physical scales: pedestrians combine root locomotion with articulated motion, whereas vehicles are rigid bodies described by kinematic state and oriented extent. Existing road-agent forecasters typically omit pedestrian articulation, while pose forecasters leave vehicle futures outside the learned rollout. We introduce PV-WM, a history-only world model over structured post-perception tracks. It recurrently advances pedestrian root motion, 15-joint articulation, and learned vehicle states within a synchronized heterogeneous state. The generated pedestrian and vehicle chunks supply the next recurrent boundary; vehicle boxes are reconstructed from predicted center and heading with observed extent, and P-V geometry is recomputed after every transition. Relative to a matched one-shot complete-state predictor, recurrent execution reduces Root ADE by 12.7% and MPJPE by 14.8%. Feedback interventions show that later predictions depend on the content, temporal order, and pedestrian identity of generated articulation. Across 824 aligned Waymo contexts, with 797 providing valid future vehicle support, PV-WM reduces Root ADE by 5.2%, MPJPE by 7.6%, P-V distance error by 11.9%, and oriented-box closest-approach error by 5.8% relative to a validation-selected Modular Specialist. The single-network model uses 57.1% fewer parameters, 96.5% lower average FLOPs per local scene, and 25.5% lower measured p95 latency. PV-WM unifies this heterogeneous future state while preserving type-specific pedestrian and vehicle dynamics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑