arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24481cs.RO

ArmnetBench v0.1:在低成本机械臂集群上对操作策略进行并行实际评估

ArmnetBench v0.1: Parallel Real-World Evaluation of Manipulation Policies on a Low-Cost Arm Farm

Praveen Selvaraj, Lorenzo Uttini, Ville Kuosmanen

AI总结:

研究针对通用机器人操作策略开发中实际评估的瓶颈问题,引入ArmnetBench v0.1基准测试,在低成本机械臂集群上对7种策略进行评估,涵盖12项任务,通过大量演示训练策略,发布核心情节数据,支持下游学习并给出排行榜比较结果。

AI中文摘要:

实际评估是通用机器人操作策略开发中的一个瓶颈。每次部署都需要物理硬件和操作员来设置、重置并评分。我们引入了ArmnetBench v0.1,这是一个在低成本SO-101单元集群上进行的、在现场轻度监督下运行的基准测试。v0.1对该机械臂集群进行端到端验证,并比较了单臂和双臂配置下12项任务中的7种策略。每个策略在每项任务上通过50次演示进行训练或微调;该基准测试包含2518次策略部署和600次参考演示。所有3118个情节都带有三元标签(成功、次优或失败)。策略部署由人工评分,而演示在构建时即设为成功。除评估外,其带有质量标签的轨迹支持下游学习,从奖励和预测世界模型到在混合质量数据上训练的策略。排行榜是在此共享预算下的初步比较。我们以LeRobot v3.0和RoboMeter格式发布3118个核心情节。

英文摘要:

Real-world evaluation is a bottleneck in developing generalist robot manipulation policies. Each rollout requires physical hardware and an operator to set up, reset, and score it. We introduce ArmnetBench v0.1, a benchmark run on a fleet of low-cost SO-101 cells under light on-site supervision. v0.1 validates this arm farm end to end and compares 7 policies across 12 tasks with both single-arm and bimanual configurations. Each policy is trained or fine-tuned on 50 demonstrations per task; the benchmark contains 2,518 policy rollouts and 600 reference demonstrations. All 3,118 episodes carry a three-way label (successful, suboptimal, or failure). Policy rollouts are human-scored, while demonstrations are successful by construction. Beyond evaluation, its quality-labelled trajectories support downstream learning, from reward and predictive world models to policies trained on mixed-quality data. The leaderboard is an initial comparison under this shared budget. We release the 3,118 core episodes in LeRobot v3.0 and RoboMeter formats.

补充信息

↑