arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14284cs.ROcs.CV

PRM-as-a-Judge 1.5:机器人过程评估工具包

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

Yuyang Liu, Yanqing Shen, Ruike Chen, Jifan Zhao, Yuxuan Tian, Yichi Zhang, Tianfeng Long, Zixuan Yin, Yipu Wang, Ziheng Qin, Wenxing Tan, Yang Shi, Mingyu Cao,… 展开作者

Yuyang Liu, Yanqing Shen, Ruike Chen, Jifan Zhao, Yuxuan Tian, Yichi Zhang, Tianfeng Long, Zixuan Yin, Yipu Wang, Ziheng Qin, Wenxing Tan, Yang Shi, Mingyu Cao, Runze Xiao, Ziqi Wang, Zhixin Yin, Shiwei Chu, Yi-Fan Zhang, Yao Mu, Yuheng Ji, Yihao Wang, Jun Yan, Zhongyuan Wang, Pengwei Wang, Xiaolong Zheng

AI总结:

该研究推出PRM-as-a-Judge 1.5工具包,新增三个评估指标并结合RoboPulse++,发布含基准与可视化工具的评估套件,呼吁建立透明可复现的机器人评估体系。

AI中文摘要:

精细的机器人评估对于理解具身模型至关重要,它超越了二元成功率和基于规则的过程评分。我们推出PRM-as-a-Judge 1.5,这是一款机器人过程评估工具包,可将 rollout 视频转化为密集的进度曲线并生成多个精细指标。PRM-as-a-Judge 1.5在1.0版本的基础上引入了三个指标,分别表征失败侧进度、回撤后恢复情况以及成功侧执行质量,帮助用户了解具身模型的能力。基于基准测试的rollout视频,我们对具身模型进行了全面评估,提供了一些精细的指标结果和关键发现。我们还推出RoboPulse++,用于评估过程奖励模型(PRM)的可靠性,为评估人员提供更准确的测试平台。此外,我们发布了一套用户友好的评估套件,包含基准测试、指标实现和可视化工具,以支持可复现的操作过程评估。我们呼吁社区重新思考机器人的评估方式,建立透明、程序化且可复现的评估,作为下一代具身智能的基础。

英文摘要:

Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.

补充信息

↑