DriftingVLA:基于逐维度时序漂移的原生单步视觉-语言-动作生成模型
DriftingVLA: Native One-Step Vision-Language-Action Generation via Per-Dimension Temporal Drifting
- University of Science and Technology of China(中国科学技术大学)
- LYNSENSE(灵犀科技)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
DriftingVLA是基于逐维度时序漂移的原生单步VLA模型,通过单步生成动作块,在多个机器人任务上优于基线模型,同时实现3.36倍的动作块生成加速。
AI中文摘要:
传统基于流的视觉-语言-动作(VLA)模型支持表达性连续动作生成,但需多步细化才能生成每个动作块,增加了在线机器人控制的延迟。为解决该问题,我们提出DriftingVLA,这是一种原生单步VLA模型,通过单个动作专家的前向传播即可生成完整动作块。与推理时需要迭代积分的流场学习不同,DriftingVLA采用分布漂移目标学习从噪声到动作块的直接映射,以实现单步部署。由于机器人动作维度具有不同的控制语义和分布特征,我们进一步提出逐维度时序漂移(PDTD),将每个动作维度的完整时序轨迹视为独立漂移单元,实现更细粒度的维度特定动作分布建模与塑造。这种逐维度分解仅应用于训练目标;共享的VLA模型仍会联合生成完整动作块,从而保留跨维度依赖关系。DriftingVLA在LIBERO上达到98.32%的成功率,在RoboTwin 2.0上达到81.09%,在6个真实世界的单臂和双臂任务上平均达到77.67%,优于所评估的多步流策略和单步VLA基线。原生单步部署还实现了动作块生成的3.36倍加速,在不牺牲控制性能的情况下消除了迭代细化。
英文摘要:
Conventional flow-based vision-language-action (VLA) models support expressive continuous action generation but rely on multi-step refinement to produce each action chunk, increasing latency in online robot control. To address this issue, we introduce DriftingVLA, a native one-step VLA that generates a complete action chunk with a single action-expert forward pass. Rather than learning a flow field that requires iterative integration at inference, DriftingVLA uses a distribution-drifting objective to learn a direct noise-to-action-chunk mapping for one-step deployment. Since robot action dimensions carry distinct control semantics and distributional characteristics, we further introduce Per-Dimension Temporal Drifting (PDTD). PDTD treats the complete temporal trajectory of each action dimension as a separate drifting unit, enabling finer-grained modeling and shaping of dimension-specific action distributions. This per-dimension decomposition applies only to the training objective; the shared VLA model still generates the complete action chunk jointly, thereby preserving cross-dimensional dependencies. DriftingVLA achieves 98.32% success on LIBERO, 81.09% on RoboTwin 2.0, and 77.67% across six real-world single- and dual-arm tasks, outperforming the evaluated multi-step flow policy and one-step VLA baselines. Native one-step deployment also delivers a 3.36-fold speedup in action-chunk generation, eliminating iterative refinement without sacrificing control performance.