HIL-UMI:将人在回路的后训练引入视觉-语言-动作模型的通用操作接口
HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface
- Peking University(北京大学)
- PrimeBot
- JD Technology(京东科技)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出HIL-UMI框架,通过人在回路的策略引导数据收集和优势细化,无需物理机器人即可对VLA模型进行后训练,在真实任务上优于监督微调。
AI中文摘要:
大规模视觉-语言-动作(VLA)模型为机器人操作提供了强大的先验知识,但将其适应到特定部署场景仍然具有挑战性。在任务特定演示数据上进行监督微调(SFT)是迈向部署的一步,但面临两个长期存在的局限:静态数据对分布外状态的覆盖有限,且标准模仿目标无法区分有进展的行为与较少有用的数据。交互式后训练可以解决这些局限,但通常需要在物理机器人上重复执行策略并进行人工干预。我们提出了HIL-UMI,一个策略引导的通用操作接口(UMI)框架,用于无需机器人的、人在回路的VLA后训练。在手持UMI演示过程中,HIL-UMI在同一观测流上查询当前策略,但不执行其预测。能量分数(Energy Score)将人类动作轨迹与策略推理进行比较,当两者差异指示分布外区域时触发数据收集。在另一个反馈回路中,低在线优势预测识别出用于细化基于进展的优势估计器的关键片段。更新后的估计器随后使用基础演示和新策略数据的平衡混合,引导优势条件的行为克隆。这种设计保留了人在回路学习的迭代性和策略感知特性,同时将数据收集与机器人部署解耦。在涵盖长时程和精细操作的四个真实世界任务上的实验表明,HIL-UMI相对于SFT取得了一致的改进,并受益于定向收集和优势细化。此外,HIL-UMI在Clean Up Table任务上优于HG-DAgger,且每帧收集时间更低,表明其为VLA后训练提供了一条可跨操作者和地点扩展的路径。
英文摘要:
Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limited coverage of out-of-distribution states, and standard imitation objectives do not distinguish progressing behavior from less useful data. Interactive post-training can address these limitations, but typically requires repeated policy execution and human intervention on a physical robot. We introduce HIL-UMI, a policy-guided Universal Manipulation Interface (UMI) framework for robot-free human-in-the-loop VLA post-training. During handheld UMI demonstrations, HIL-UMI queries the current policy on the same observation stream without executing its predictions. The Energy Score compares the human action trajectory with policy inference and triggers collection when their discrepancy indicates an out-of-distribution region. In a separate feedback loop, low online advantage predictions identify essential segments for refining a progress-based advantage estimator. The updated estimator then guides advantage-conditioned behavioral cloning using a balanced mixture of base demonstrations and new policy data. This design preserves the iterative and policy-aware nature of human-in-the-loop learning while decoupling data collection from robot deployment. Experiments on four real-world tasks spanning long-horizon and precise manipulation show that HIL-UMI achieves consistent improvement over SFT and benefits from both targeted collection and advantage refinement. Moreover, HIL-UMI outperforms HG-DAgger on Clean Up Table with lower per-frame collection time, suggesting a scalable path for VLA post-training across operators and locations.