arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VGGTWorld-VLA:面向自动驾驶的意图条件3D世界演化模型

VGGTWorld-VLA: Intent-Conditioned 3D World Evolution for Autonomous Driving

Zhaoyang Liu, Kun Jiang, Ziying Song, Diange Yang

arXiv 2610.11161首次发表:更新:

发表机构

Tsinghua University; Nanyang Technological University(清华大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有自动驾驶3D世界模型对驾驶意图条件依赖弱的问题,提出VGGTWorld-VLA模型,通过动作-语义条件机制和几何-语言-动作桥实现可控未来几何预测,在NAVSIM上验证了方法的有效性。

AI 中文摘要

VGGT为以几何为核心的世界模型提供了坚实基础,可从视觉观测中恢复统一的3D场景几何。尽管近期扩展模型支持时序3D预测,但其未来演化对驾驶意图和动作的条件依赖较弱,限制了其对不同动作依赖的未来场景的建模能力。我们提出VGGTWorld-VLA,这是VGGT-World的意图条件扩展模型,用于自动驾驶中可控的3D世界演化。首先,我们引入动作-语义条件机制,将互补的驾驶语义和 ego-motion(自运动)表征注入未来token流,使同一观测场景在不同自车动作下能生成不同的未来几何预测。其次,我们构建几何-语言-动作桥,适配历史几何、VLA语义特征以及操纵和轨迹表征,用于未来几何预测的联合条件。我们在NAVSIM数据集上评估未来几何预测性能,同时通过条件消融研究检验语义和动作信息的贡献。与基线相比,我们的方法展现出具有竞争力的几何预测性能;消融研究进一步验证了语义和动作条件的有效性。这些结果表明,语义和动作条件在自动驾驶中基于VGGT的可控世界预测方面具有潜力。

英文摘要

VGGT provides a strong foundation for geometry-centric world models by recovering unified 3D scene geometry from visual observations. Although recent extensions enable temporal 3D prediction, their future evolution remains weakly conditioned on driving intentions and actions, limiting their ability to model alternative action-dependent futures. We propose VGGTWorld-VLA, an intention-conditioned extension of VGGT-World for controllable 3D world evolution in autonomous driving. First, we introduce an action--semantic conditioning mechanism that injects complementary driving semantics and ego-motion representations into the future-token stream, enabling different future geometry predictions for the same observed scene under alternative ego actions. Second, we develop a geometry--language--action bridge that adapts historical geometry, VLA semantic features, and maneuver and trajectory representations for joint conditioning of future geometry prediction. We evaluate future geometry prediction on NAVSIM, while conditioning ablations further examine the contributions of semantic and action information. Compared with the baseline, our method demonstrates competitive geometry prediction performance. Ablation studies further support the effectiveness of semantic and action conditioning. These results demonstrate the potential of semantic and action conditioning for controllable VGGT-based world prediction in autonomous driving.

Comments20 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑