arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAF-VLA:面向端到端自动驾驶的未来表征对齐

RAF-VLA: Representation Alignment with the Future for End-to-End Autonomous Driving

Dogun Kim, Yongjae Lee, Joonhee Lim, Yeina Lee, Junhyeok Park, Moogeun Park, Dongsuk Kum

arXiv 2609.17728首次发表:更新:

发表机构

Korea Advanced Institute of Science & Technology (KAIST)(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对世界建模VLA在自动驾驶中因显式未来生成带来的训练负担和推理延迟问题,提出RAF-VLA框架,通过未来表征对齐的监督微调直接引导内部表征,在NAVSIM基准上以更少训练样本达到竞争性能,且训练开销仅3.8%、推理开销仅1毫秒。

AI 中文摘要

近期用于自动驾驶的视觉-语言-动作(VLA)模型通过预测未来驾驶场景以及驾驶动作来融入世界建模,展现出强大的规划性能。未来驾驶场景被用作密集监督,促使策略学习对规划有用的丰富内部表征。然而,这些世界建模VLA依赖显式的未来生成来学习此类表征,从而引入了两个关键限制:额外的训练负担和推理延迟。为解决这些限制,我们提出了RAF-VLA(面向未来的表征对齐),一种基于VLA的自动驾驶框架,通过未来帧表征的直接引导来塑造与规划相关的内部表征。RAF-VLA采用未来对齐的监督微调,其中一项简单的正则化在策略学习驾驶动作的同时,将其隐藏状态与从预训练世界编码器获得的未来帧表征对齐。这种简单的对齐使RAF-VLA能够避免与未来生成相关的训练负担和推理延迟。在NAVSIM基准上的大量实验表明,RAF-VLA在显著减少训练样本的情况下,达到了与最先进的VLA规划器相当的规划性能。此外,RAF-VLA仅产生3.8%的训练开销和可忽略不计的1毫秒推理开销。

英文摘要

Recent Vision-Language-Action (VLA) models for autonomous driving have incorporated world modeling by predicting future driving scenes alongside driving actions, demonstrating strong planning performance. Future driving scenes are utilized as dense supervision, encouraging the policy to learn rich internal representations useful for planning. However, these World-Modeling VLAs rely on explicit future generation to learn such representations, thereby introducing two key limitations: additional training burden and inference latency. To address these limitations, we propose RAF-VLA (Representation Alignment with the Future), a VLA-based autonomous driving framework that shapes planning-relevant internal representations through direct guidance from future-frame representations. RAF-VLA employs Future-Aligned Supervised Fine-Tuning, in which a straightforward regularization aligns the policy's hidden states with future-frame representations obtained from a pretrained world encoder while learning driving actions. This simple alignment allows RAF-VLA to avoid the training burden and inference latency associated with future generation. Extensive experiments on the NAVSIM benchmark show that RAF-VLA achieves competitive planning performance against state-of-the-art VLA planners with substantially fewer training samples seen. Moreover, RAF-VLA incurs only 3.8% training overhead and a negligible 1 ms inference overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑