R2S-Eval:基于视觉语言模型的现实到仿真校准的机器人评估方法
R2S-Eval: Robot Evaluation with Real-to-Sim Calibration via Vision-Language Models
另 1 家 · 查看机构详情
- Sharpa(夏普)
- Nanjing University(南京大学)
- Tongji University(同济大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
R2S-Eval结合现实到仿真校准与VLM偏好评估,实现机器人操纵策略的自动化、稳定且感知质量的评估,减少硬件操作工作量并揭示二元成功标签未捕捉的行为差异。
中文摘要 AI 辅助
随着通用模型,特别是视觉语言动作(VLA)模型被部署到物理机器人上,评估机器人操纵策略变得愈发重要。然而,传统的现实世界评估仍然劳动密集、不稳定且信息不足:它需要重复的硬件试验、手动场景重置和持续的操作员监控,可能在重复评估中产生不同的策略排名,且主要依赖成功率指标,该指标对执行质量的信息提供有限。相比之下,人类通过观察和比较完整行为来评估机器人性能,而非仅依赖二元成功结果。为此,我们提出R2S-Eval,一种结合现实到仿真校准与视觉语言模型(VLM)偏好评估的评估流程。现实到仿真组件在校准至现实世界评估设置的模拟器中高效生成滚动视频,从而减少对重复硬件试验的需求。VLM评估器评估滚动视频的执行质量并生成成对偏好,随后将其聚合成策略排名。我们进一步引入一项协议,以评估所提出的评估流程是否产生经验证的策略结论,同时缓解传统现实世界评估的关键挑战。在仿真和现实世界设置中的实验表明,R2S-Eval产生可靠且稳定的策略结论,与人类偏好达成一致,大幅减少重复的硬件操作工作量,并揭示二元成功标签未捕捉到的行为质量差异。总体而言,R2S-Eval将机器人评估从手动成功计数推进到对机器人行为的自动化、统计稳定且感知质量的评估。项目页面:this https URL。
英文摘要
Evaluating robot manipulation policies is becoming increasingly important as generalist models, particularly vision-language-action (VLA) models, are deployed on physical robots. However, conventional real-world evaluation remains labor-intensive, unstable, and insufficiently informative. It requires repeated hardware trials, manual scene resets, and continuous operator monitoring, may produce different policy rankings across repeated evaluations, and primarily relies on success-rate metrics that provide limited information about execution quality. In contrast, humans assess robot performance by observing and comparing complete behaviors rather than relying solely on binary success outcomes. To this end, we propose R2S-Eval, an evaluation pipeline that combines real-to-sim calibration with vision-language model (VLM) preference evaluation. The real-to-sim component efficiently generates rollout videos in a simulator calibrated to the real-world evaluation setting, thereby reducing the need for repeated hardware trials. The VLM evaluator assesses the execution quality of rollout videos and produces pairwise preferences, which are subsequently aggregated into policy rankings. We further introduce a protocol to assess whether the proposed evaluation pipeline yields validated policy conclusions while mitigating the key challenges of conventional real-world evaluation. Experiments in both simulation and real-world settings demonstrate that R2S-Eval produces reliable and stable policy conclusions, achieves agreement with human preferences, substantially reduces repeated hardware-operation effort, and reveals behavior-quality differences that are not captured by binary success labels. In general, R2S-Eval advances robot evaluation from manual success counting toward automated, statistically stable, and quality-aware evaluation of robot behavior. Project page: https://r2s-eval.github.io.