AI 中文总结
Track4Action框架将以世界为中心的冻结3D跟踪器提炼为视觉-语言-动作策略,在多个机器人任务基准上显著提升性能,验证了动作对齐的3D跟踪器特征作为监督的有效性。
AI 中文摘要
动作标签会告知视觉-语言-动作(VLA)策略应模仿哪些机器人指令,但不会说明这些指令如何改变3D世界。对齐的演示片段包含了缺失的监督信息,因为其K帧过渡记录了对应K个动作过程中产生的几何形状、运动、可见性和相机变化。我们提出Track4Action,这是一个将以世界为中心的冻结3D跟踪器中实现的过渡提炼为当前观测VLA策略的框架。训练期间,Track4World将片段Vₜ:ₜ₊ₖ编码为池化跟踪器特征;可学习的跟踪查询从当前VLA隐藏状态推断该特征,在共享空间中对其进行匹配,并通过特征门控调节流匹配动作头。跟踪器特征仅定义对齐目标,因此部署时既不使用片段也不使用跟踪器。Track4Action在零样本LIBERO-Plus上达到82.3%的性能,比无对齐变体提升7.6个百分点,比LaMP提升3.0个百分点;在干净和随机RoboTwin 2.0分割上分别达到80.44%和81.48%;在四项物理双臂任务上平均成功率为67.5%,比无对齐变体高出25.0个百分点。模拟和物理任务上的增益表明,动作对齐的3D跟踪器特征是无跟踪器VLA部署的特权监督。我们的项目页面可在该https URL获取。
英文摘要
Action labels tell a vision-language-action (VLA) policy which robot commands to imitate, but not how those commands change the 3D world. The aligned demonstration clip contains this missing supervision because its $K$ frame transitions record the geometry, motion, visibility, and camera change produced during the corresponding $K$ actions. We introduce Track4Action, a framework that distills this realized transition from a frozen world-centric 3D tracker into a current-observation VLA policy. During training, Track4World encodes the clip $V_{t:t+K}$ into a pooled tracker feature. Learnable track queries infer this feature from current VLA hidden states, match it in a shared space, and condition a flow-matching action head through a feature-wise gate. The tracker feature only defines the alignment target, so neither the clip nor the tracker is used at deployment. Track4Action reaches 82.3% on zero-shot LIBERO-Plus, improving the alignment-free variant by 7.6 points and LaMP by 3.0 points. It obtains 80.44% and 81.48% on the clean and randomized RoboTwin 2.0 splits, and 67.5% average success across four physical bimanual tasks, 25.0 points above the alignment-free variant. The gains across simulation and physical tasks support action-aligned 3D tracker features as privileged supervision for tracker-free VLA deployment. Our project page is available at https://wing0night.github.io/track4action-project-page.