arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

弥合实验室与商店之间的差距:一种用于零售类人机器人的数据高效训练后及经验驱动学习的VLA框架

Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids

Roger Sala Sisó, Tiago Silvério, Jakob Sand, Tran Nguyen Le

arXiv 2607.20345首次发表:更新:

发表机构

HIVE Robots; Technical University of Denmark(蜂巢机器人公司; 丹麦技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对VLA类人机器人弥合实验室与实际应用差距的问题,提出DEED系统级方法,含数据高效训练后管道、经验驱动细化研究及潜在空间分析工具,经实验证明精心设计数据和训练后处理可解决该问题。

AI 中文摘要

弥合基准性能与可靠的实际操作之间的差距仍然是视觉语言动作(VLA)类人机器人面临的核心挑战,这类机器人必须应对执行错误、分布变化和环境变异性。本文提出了DEED(数据高效训练后及经验驱动学习),这是一种在超市芯片补货任务中使用宇树G1-Edu类人机器人和GR00T N1.6基础模型进行评估的系统级方法。DEED包括三个关键组件:(1)一个具有控制频率对齐、数据管理、任务相关视觉突出显示和降低VLA依赖性的数据高效训练后管道;(2)通过基于文本的优势前缀和视觉语言价值函数从RECAP改编而来的经验驱动细化的实际研究;(3)用于研究分布内和分布外行为的潜在空间分析工具。我们的结果表明,弥合实验室与商店之间的差距主要是系统集成挑战而非架构挑战:精心的数据设计和有针对性的训练后处理可以仅使用单个GPU将在简单微调下失败的策略转变为一个能胜任实际操作的系统。

英文摘要

Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA) humanoid robots, which must handle execution errors, distribution shifts, and environmental variability. This paper presents DEED (Data-Efficient Post-Training and Experience-Driven Learning), a systems-level approach evaluated on a supermarket chip-restocking task using a Unitree G1-Edu humanoid robot and the GR00T N1.6 foundation model. DEED comprises three key components: (1) a data-efficient post-training pipeline with control-frequency alignment, data curation, task-relevant visual highlighting, and reduced VLA dependence; (2) a real-world study of experience-driven refinement, adapted from RECAP via a text-based advantage prefix and a vision-language value function; and (3) a latent-space analysis tool for studying in- and out-of-distribution behavior. Our results suggest that bridging the lab-to-store gap is primarily a systems integration challenge rather than an architectural one: careful data design and targeted post-training can transform a policy that fails under naive fine-tuning into a competent real-world system using only a single GPU.

Comments8 pages. This work has been submitted to the IEEE for possible publication

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑