发表机构
Automotive New Technology Research Institute, BYD Company Limited(比亚迪汽车新技术研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对自动驾驶中视觉-语言-动作模型的不足,提出HyWorldVLA框架,统一像素级监督与潜在表征学习。预训练阶段预测视频潜在并重建帧,微调阶段预测潜在特征生成轨迹,实验表明其性能优于基线,还建立了世界模型噪声鲁棒性评估新基准。
AI 中文摘要
增强世界建模的视觉-语言-动作(VLA)模型是端到端自动驾驶的一个有前景的范式。像素级未来预测虽能进行细粒度时空推理,但在嘈杂驾驶场景中鲁棒性不足;基于潜在的世界模型能缓解敏感性问题,却存在可解释性有限和表征退化的问题。为调和这种权衡,我们提出HyWorldVLA,一个统一像素级监督和潜在表征学习的混合世界-VLA框架。在预训练阶段,HyWorldVLA预测由预训练视频VAE编码的视频潜在,同时重建视频帧以提供精确的像素级基础。在随后的联合微调阶段,模型专门预测潜在特征,输入动作专家生成轨迹。在NAVSIM v1和v2基准上的大量实验表明,HyWorldVLA显著优于基于像素和基于潜在的世界模型基线。值得注意的是,我们首次对自动驾驶中世界模型噪声鲁棒性进行了全面的定性和定量分析,为评估未来架构建立了新基准。
英文摘要
Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level future prediction enables fine-grained spatiotemporal reasoning, it compromises robustness in noisy driving scenarios. Conversely, latent-based world models alleviate this sensitivity but often incur limited interpretability and representational degradation due to absent pixel-level grounding. To reconcile this trade-off, we propose HyWorldVLA, a hybrid world-VLA framework that unifies pixel-level supervision and latent representation learning. In the pre-training stage, HyWorldVLA predicts video latents encoded by a pre-trained video VAE, while simultaneously reconstructing video frames to provide precise pixel-level grounding. During the subsequent co-fine-tuning phase, the model exclusively predicts latent features, which are fed into an action expert to generate trajectories. Extensive experiments on NAVSIM v1 and v2 benchmarks demonstrate that HyWorldVLA significantly outperforms both pixel-based and latent-based world model baselines. Notably, we present the first comprehensive qualitative and quantitative analysis of world model noise robustness in autonomous driving, establishing a new benchmark for evaluating future architectures.
Comments20 pages with 13 figures