发表机构
DAMO Academy, Alibaba Group; Hupan Lab(达摩院(阿里巴巴集团); 湖畔实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出开源机器人价值基础模型 RynnValue,以时间距离为监督目标,在多维度超越现有方法,可提升机器人策略的真实世界任务成功率并实现零样本泛化。
AI 中文摘要
通用奖励模型正日益成为扩展机器人学习的瓶颈,但从大规模异构语料库中学习价值相关能力的方法仍未得到充分探索。现有方法将监督与任务内部锚点(如偏好或归一化进度)绑定,这些锚点均无法跨 embodiment 和数据源实现干净迁移。我们推出 RynnValue,一款面向机器人操纵的开源价值基础模型,它用时间距离(从观测到语言指定目标的定向后续成本)替代上述锚点。由于时间距离标签可直接从时间戳推导,RynnValue 可扩展至超 7000 小时、约 300 万条指令条件片段,无需偏好或进度标注。为使时间价值学习在大规模场景下可靠,我们结合随机时间采样、时间顺序打乱和价值隔离注意力,抑制会导致预测对失败和倒退不敏感的捷径。在无偏好标签的情况下,RynnValue 在 RBM-EVAL-OOD 上取得平均 Kendall's tau_a 为 0.675 的成绩,超过完全偏好监督的现有最优方法(0.655),是仅基于进度的对应模型(0.292)的两倍多,同时可零样本泛化至未见过的任务、 embodiment 和视角。通过基于势能的塑形转换为密集奖励后,它将真实世界策略的在线成功率从 52.5% 提升至 72.5%,离线成功率从 63.8% 提升至 82.5%。这些结果确立时间距离为通用机器人策略的可扩展监督目标和实用奖励接口。
英文摘要
General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-source value foundation model for robotic manipulation that replaces these anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal. Because temporal-distance labels can be derived directly from timestamps, RynnValue scales to over 7,000 hours and roughly 3M instruction-conditioned clips without preference or progress annotations. To make temporal-value learning reliable at scale, we combine random temporal sampling, temporal-order shuffling, and value-isolation attention, suppressing shortcuts that would leave predictions insensitive to failures and regressions. Trained without preference labels, RynnValue attains an average Kendall's $τ_a$ of 0.704 on RBM-EVAL-OOD, surpassing the fully preference-supervised state of the art (0.655) and more than doubling a progress-only counterpart (0.292), while generalizing zero-shot to unseen tasks, embodiments, and viewpoints. As a zero-shot reward model, RynnValue serves a range of downstream applications. Converted into dense rewards via potential-based shaping, it raises real-world policy success from 52.5% to 72.5% online and from 63.8% to 82.5% offline; used for data filtering, it improves multi-task behavior cloning success from 35.0% to 42.5%; and applied as inference-time value guidance, it lifts a frozen policy's success from 67.5% to 80.0%. These results establish temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.
Comments32 pages, 7 figures