发表机构
Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉-语言-动作模型在有限演示下适配性能下降的问题,提出分层技能检索框架,结合子任务语言检索与行为特征重排序,在LIBERO基准及真实机器人任务上提升平均成功率10.3%和21.3%。
AI 中文摘要
尽管在大规模机器人数据集上预训练的视觉-语言-动作(Vision-Language-Action,VLA)模型为机器人操作提供了强大的基础,但当适配到仅有有限任务特定演示的新任务时,其性能可能会下降。检索为复用现有演示提供了一种实用的数据高效适配方式,然而现有方法通常依赖视觉相似度、状态-动作表示或任务级语言匹配,这些方法可能忽略长程操作任务的分层结构——在这类任务中,完整任务匹配较为罕见,但可复用的技能往往十分丰富。为应对这一挑战,我们提出分层技能检索(Hierarchical Skill Retrieval,HSR),这是一种用于视觉-语言-动作模型数据高效适配的检索框架。具体而言,HSR首先将目标任务分解为候选技能序列,它基于语义合理性和从先验数据集估计的技能可靠性评估每个计划,随后将选定的分解结果用于混合检索,该过程将子任务级语言检索与行为特征重排序相结合,以识别既语义相关又与目标任务兼容的演示。最后,我们通过两阶段预训练与微调流程适配策略,该流程将通用技能获取与任务特定适配分离开来。在LIBERO基准及多个真实世界机器人操作任务上的实验表明,HSR相比最强基线分别将平均成功率提高了10.3%和21.3%,这些结果证明了结构化技能级检索在视觉-语言-动作模型数据高效适配中的有效性。视频和代码可在此https URL获取。
英文摘要
While Vision-Language-Action (VLA) models pretrained on large-scale robot datasets provide a strong foundation for robot manipulation, their performance can degrade when adapted to new tasks with limited task-specific demonstrations. Retrieval offers a practical way to reuse existing demonstrations for data-efficient adaptation, but existing methods often rely on visual similarity, state-action representations, or task-level language matching. These approaches may overlook the hierarchical structure of long-horizon manipulation tasks, where complete task matches are rare but reusable skills are often abundant. To address this challenge, we propose Hierarchical Skill Retrieval (HSR), a retrieval framework for data-efficient VLA adaptation. Specifically, HSR first decomposes a target task into candidate skill sequences. It evaluates each plan based on both semantic plausibility and skill reliability estimated from the prior dataset. The selected decomposition is then used for hybrid retrieval. This combines subtask-level language retrieval with behavior-feature reranking to identify demonstrations that are both semantically relevant and compatible with the target task. Finally, we adapt the policy through a two-stage pretraining and finetuning pipeline, which separates general skill acquisition from task-specific adaptation. Experiments on the LIBERO benchmark and several real-world robot manipulation tasks show that HSR improves the average success rate by 10.3% and 21.3% over the strongest baseline, respectively. These results demonstrate the effectiveness of structured skill-level retrieval for data-efficient VLA adaptation. Videos and code are available at https://hoar012.github.io/HSR-Project.
CommentsProject Page: https://hoar012.github.io/HSR-Project