arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReWeight:利用人类数据进行VLA后训练——通过演示检索与样本加权

ReWeight: Leveraging Human Data for VLA Post-Training via Demonstration Retrieval and Sample Weighting

Chenwei Wang, Dianye Huang, Match W. L. Ko, Chenjia Bai, Zhongliang Jiang

arXiv 2609.13851首次发表:更新:

发表机构

The University of Hong Kong; Institute of Artificial Intelligence (TeleAI), China Telecom; Shenzhen Research Institute of Northwestern Polytechnical University(香港大学; 中国电信人工智能研究院(TeleAI); 西北工业大学深圳研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ReWeight通过演示检索与样本加权,将人类演示数据有效融入VLA模型后训练,提升跨具身任务成功率,模拟中从39%提至57%,真实任务达68.8%。

AI 中文摘要

针对特定机器人和任务的后训练视觉-语言-动作(VLA)模型需要领域内的演示数据,然而收集多样化的机器人数据成本高昂。以自我为中心的人类演示提供了一种可扩展的替代方案,但直接将人类数据与机器人数据混合可能引入跨具身差异并降低策略性能。为解决这一挑战,我们提出了ReWeight,一个通过演示级检索和样本级加权将人类数据纳入VLA后训练的框架。ReWeight学习一种跨具身视觉运动表征,该表征结合视觉观察与未来动作,以衡量人类演示与机器人演示之间的行为相似性。基于最优传输,它检索与目标机器人数据相关的人类演示,并为跨具身差异较小的样本分配更大的权重。我们使用π0.5在八个模拟任务和四个真实世界任务中,在干净和随机设置下评估了ReWeight。在模拟中,ReWeight将后训练π0.5的平均成功率从仅使用机器人数据时的39%和随机混合人类-机器人数据时的44%提升至57%。在物理实验设置中,它达到了68.8%的平均成功率,分别比基线高出28.8%和13.8%。总体而言,ReWeight提供了一种有效的范式,将丰富的以自我为中心的人类经验转化为可迁移的机器人学习监督。(项目网页:此https URL)

英文摘要

Post-training vision-language-action (VLA) models for specific robots and tasks requires in-domain demonstrations, yet collecting diverse robot data is costly. Egocentric human demonstrations provide a scalable alternative, but directly mixing human and robot data can introduce cross-embodiment discrepancies and degrade policy performance. To address this challenge, we introduce ReWeight, a framework that incorporates human data into VLA post-training through demonstration-level retrieval and sample-level weighting. ReWeight learns a cross-embodiment visuomotor representation that combines visual observations with future actions to measure behavioral similarity between human and robot demonstrations. Based on optimal transport, it retrieves human demonstrations relevant to the target robot data and assigns larger weights to samples with smaller cross-embodiment discrepancies. We evaluate ReWeight using $π_{0.5}$ across eight simulation tasks and four real-world tasks under both clean and randomized settings. In simulation, ReWeight improves the average success rate of post-trained $π_{0.5}$ from 39% with only robot data and 44% with randomly mixed human-robot data to 57%. In the physical experimental setting, it achieves an average success rate of 68.8%, outperforming the baselines by 28.8% and 13.8%, respectively. Overall, ReWeight provides an effective paradigm for transforming abundant egocentric human experience into transferable supervision for robot learning. (Project webpage: https://reweight-vla.github.io/)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑