发表机构
XPeng Inc(小鹏汽车)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出XCoT-VLA,用可执行CoT令牌替代自然语言CoT,结合Reason FFN与Control FFN实现轨迹生成,在自动驾驶任务中降低了误差并满足实时规划要求。
AI 中文摘要
视觉-语言-动作(VLA)模型可将场景理解、语义推理与自主驾驶的轨迹生成关联起来。然而,冗长的自然语言思维链(CoT)不适合实时控制,因为它开放性强、解码成本高,且难以作为面向动作的表示进行优化。我们提出XCoT-VLA,它用从自动构建的推理-动作监督中学习到的紧凑可执行CoT令牌替代描述性理由。记录的轨迹提供动作证据,而场景上下文提供因果语义。预测的XCoT序列保持在上下文中,并通过共享多模态自注意力条件化固定轨迹查询。确定性令牌-函数路由将推理前馈网络(Reason FFN)应用于XCoT令牌,将控制前馈网络(Control FFN)应用于轨迹查询以进行流匹配轨迹生成。我们进一步引入XCoT策略优化(XCPO)作为同一可执行令牌空间中的可选优化扩展。XCoT-VLA在通用分布集上将纵向平均位移误差(ADE)从1.645降至1.323,在变道场景上将横向最终位移误差(FDE)从1.616降至0.648。通过仅用2-6个可执行XCoT令牌表示面向驾驶的推理,我们的方法大幅降低了自回归推理开销,并保持在实时规划预算内。这些结果表明,面向驾驶的推理可以是紧凑、可执行且直接与轨迹生成相连的。
英文摘要
Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verbose natural-language Chain-of-Thought (CoT) is poorly suited to real-time control because it is open-ended, costly to decode, and difficult to optimize as an action-facing representation. We propose XCoT-VLA, which replaces descriptive rationales with compact executable CoT tokens learned from automatically constructed Reason-Action supervision. Logged trajectories provide action evidence, while scene context supplies causal semantics. The predicted XCoT sequence remains in context and conditions fixed trajectory queries through shared multimodal self-attention. Deterministic token-function routing applies the Reason FFN to XCoT tokens and the Control FFN to trajectory queries for flow-matching trajectory generation. We further introduce XCoT Policy Optimization (XCPO) as an optional refinement extension in the same executable token space. XCoT-VLA reduces longitudinal ADE from 1.645 to 1.323 on a general-distribution set and lateral FDE from 1.616 to 0.648 in lane-change scenarios. By representing driving-oriented reasoning with only 2-6 executable XCoT tokens, our method substantially reduces autoregressive reasoning overhead and remains within the real-time planning budget. These results demonstrate that driving-oriented reasoning can be compact, executable, and directly connected to trajectory generation.