AI 中文总结
研究针对机器人学习中时间衍生进度标签不准确问题,提出无监督机器人价值校正方法UR-VC,利用演示数据中相似状态不同时出现的规律校正标签,在真实双手布料操作任务中取得积极效果,还能用于构建优势标签。
AI 中文摘要
现代机器人学习系统越来越依赖密集的进度或价值信号来评估中间状态、指导策略学习和检测任务完成,因此这些信号的质量至关重要。由于此类密集标签很少大规模可用,演示中的归一化时间常被用作可扩展的替代:后期帧被视为更高进度。但这种时间衍生标签只是物理任务进度的噪声代理。在丰富接触的操作中,机器人可能取得进展后又因滑倒、抓握失败或部分撤销而失去进展,而时间衍生标签仍单调增加。我们引入无监督机器人价值校正(UR-VC),一种用于校正时间衍生进度标签的离线、无需训练的方法。UR-VC利用演示数据中的简单规律:相似状态常在不同情节中不同时间戳重复出现。它从其他情节中检索相似状态并聚合其时间衍生标签以获得校正后的进度估计,无需人工进度标签、奖励注释或额外价值模型。我们在真实的双手布料平整和折叠数据上评估UR-VC,这是一个具有可见中间进度的长时程可变形物体操作任务。校正后的标签捕捉了归一化时间无法表示的局部回归和非均匀进度,同时保留了整体任务趋势。我们还进一步使用校正后的信号为VLA训练构建优势标签。在匹配的数据、模型和训练设置下,UR-VC在真实机器人任务成功方面呈现积极趋势。
英文摘要
Modern robot learning systems increasingly rely on dense progress or value signals to evaluate intermediate states, guide policy learning, and detect task completion, making the quality of these signals critical. Since such dense labels are rarely available at scale, normalized time within a demonstration is often used as a scalable substitute: later frames are treated as higher progress. However, this time-derived label is only a noisy proxy for physical task progress. In contact-rich manipulation, a robot may make progress and then lose it through slips, failed grasps, or partial undoing, while the time-derived label continues to increase monotonically. We introduce Unsupervised Robotic Value Correction (UR-VC), an offline, training-free method for correcting time-derived progress labels. UR-VC exploits a simple regularity in demonstration data: similar states often recur across different episodes, but at different timestamps. Instead of trusting the timestamp from a single trajectory, UR-VC retrieves similar states from other episodes and aggregates their time-derived labels to obtain a corrected progress estimate. UR-VC requires no manual progress labels, reward annotations, or additional value model. We evaluate UR-VC on real bimanual cloth flatten-and-fold data, a long-horizon deformable-object manipulation task with visible intermediate progress. The corrected labels capture local regressions and non-uniform progress that normalized time cannot represent, while preserving the overall task trend. We further use the corrected signal to construct advantage labels for VLA training, following recent advantage-conditioned policy learning. UR-VC shows a positive trend in real-robot task success under matched data, model, and training settings.