arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

遗留数据何时开始发挥作用?跨配置机器人学习中的涌现转移

When Does Legacy Data Start to Help? Emergent Transfer in Cross-Configuration Robot Learning

Tao Wang, Hudson Hou, Yingdong Hu, Yufeng Liu, Qinghai Li, Yingjie Jiang, Yingzhi Wang, Cheng Ma, Richard Wang, Yang Gao

arXiv 2607.25593首次发表:更新:

AI 中文总结

研究跨配置机器人学习中遗留数据何时有益,发现存在类似顿悟的三阶段转变,即低能力时遗留数据无效,超阈值后收益剧增,高能力时收益递减,基于梯度对齐等给出理论解释和阶段感知规则并验证。

AI 中文摘要

机器人硬件随时间演变,但示范数据往往与特定传感器和执行器配置相关联。这引发了一个实际且未被充分探索的问题:遗留数据何时开始对升级后的机器人有益?我们在一个轮式人形平台上跨两代硬件研究此问题,相机和抓手都改变而整体形态不变。与更多跨配置数据总是有益的常见假设相反,我们观察到类似顿悟的转变:在升级配置获得最低任务能力水平之前,遗留数据无效,之后联合训练收益急剧上升,在接近饱和时下降。我们假设这种任务依赖的转变由转移阈值控制并描述了由此产生的三阶段模式。在实际机器人操作任务中,我们观察到了所有三个阶段:低能力时无显著益处(从10.0%到10.0%),超过阈值后大幅提升(插花任务从23.3%到86.7%),高能力时收益递减(笔插入任务从85.0%到93.3%)。我们基于梯度对齐和残余策略不确定性提供了理论解释,并得出了一个阶段感知规则,用于决定何时收集更多新硬件数据以及何时重用遗留示范。我们还在移动双臂浇水任务上验证了这种三阶段模式,结果与我们的预测一致。

英文摘要

Robotic hardware evolves over time, but demonstration data is often tied to a specific sensor and actuator configuration. This raises a practical and underexplored question: when does legacy data begin to benefit an upgraded robot? We study this question on a wheeled humanoid platform across two hardware generations, where both the camera and gripper are changed while the overall morphology remains fixed. Contrary to the common assumption that more cross-configuration data is always helpful, we observe a grokking-like transition: legacy data remains ineffective until the upgraded configuration acquires a minimum level of task competence, after which co-training gains rise sharply before diminishing near saturation. We hypothesize that this task-dependent transition is governed by a transfer threshold and characterize the resulting three-phase pattern. Across real-robot manipulation tasks, we observe all three phases: no measurable benefit at low competence ($10.0\% \rightarrow 10.0\%$), a sharp gain after crossing the threshold ($23.3\% \rightarrow 86.7\%$ on flower insertion), and diminishing returns at high competence ($85.0\% \rightarrow 93.3\%$ on pen insertion). We provide a theoretical account based on gradient alignment and residual policy uncertainty, and derive a phase-aware rule for deciding when to collect more new-hardware data and when to reuse legacy demonstrations. We further validate this three-phase pattern on a mobile dual-arm watering task, with results consistent with our predictions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑