面向高自由度灵巧操作的后训练VLA方法
Towards High-DoF Dexterous Manipulation through VLA Post-Training
浏览论文内容
中文总结 AI 辅助
提出四步后训练流程(动作编解码器、监督微调、DAgger、残差强化学习),解决高自由度灵巧手VLA部署难题,在五项真实任务上均达100%成功率。
中文摘要 AI 辅助
模仿学习的视觉-语言-动作(VLA)基础模型通过跨任务和跨本体扩展机器人数据来获得广泛的操作能力,但在特定下游任务和硬件平台上可靠部署仍需后训练。灵巧手使这种适应尤为困难:其广泛的行为技能库和高自由度产生了庞大且结构化的动作空间。三个核心障碍是:开源VLA原生不提供高自由度手的动作接口;在人工门控DAgger接管过程中的手势不匹配会造成命令不连续并污染纠正轨迹;在原始关节空间中的强化学习样本效率低下。我们提出一个统一的四步后训练流程,包括学习式时间手部动作编解码器、监督微调、DAgger和真实世界残差强化学习。该编解码器将预训练VLA适配到绝对灵巧手命令。缓冲回滚、姿态对齐和平滑命令混合实现了连续且任务相关的DAgger纠正,而潜在残差强化学习将探索限制在编解码器捕获的协调手部运动内。我们在五个多样化的真实世界任务上评估该流程,涵盖双臂转移、手内重定向和工具使用。在报告的后训练预算内,所得策略在每项任务20次试验中均达到100%的成功率。这些结果为将VLA基础模型适配到可靠的真实世界灵巧操作提供了实用路径。
英文摘要
Imitation-learned vision--language--action (VLA) foundation models acquire broad manipulation capabilities by scaling robot data across tasks and embodiments, but reliable deployment on a specific downstream task and hardware platform still requires post-training. Dexterous hands make this adaptation particularly difficult: their broad behavioural repertoire and high degree of freedom create a large and structured action space. Three obstacles are central: open-source VLAs do not natively provide an action interface for high-DoF hands; gesture mismatch during human-gated DAgger takeover creates command discontinuities and contaminates corrective trajectories; and reinforcement learning in the raw joint space is sample-inefficient. We present a unified four-step post-training pipeline comprising a learned temporal hand-action codec, supervised fine-tuning, DAgger, and real-world residual reinforcement learning. The codec adapts a pretrained VLA to absolute dexterous-hand commands. Buffered rollback, pose alignment, and smooth command blending enable continuous, task-relevant DAgger corrections, while latent residual RL confines exploration to coordinated hand motions captured by the codec. We evaluate the pipeline on five diverse real-world tasks spanning bimanual transfer, in-hand reorientation, and tool use. Within the reported post-training budgets, the resulting policies achieve 100\% success on every evaluated task over 20 trials per task. These results provide a practical path for adapting VLA foundation models to reliable real-world dexterous manipulation.
发表机构
- Wuji Technology(无极科技)
- School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。