基于因子的通用机器人策略的主动真实世界评估
Active Real-World Factor-Based Evaluation for Generalist Robot Policies
浏览论文内容
中文总结 AI 辅助
研究通用机器人策略评估难题,提出主动评估框架,将策略评估当作顺序实验设计问题,通过拟合概率替代模型、自适应选评估配置,高效表征策略行为并识别易失败区域,相比随机测试节省大量试验次数。
中文摘要 AI 辅助
在大型多样数据集上训练的通用机器人操作策略在广泛任务中显示出巨大潜力。然而,严格评估这些策略仍是根本挑战。现实世界性能取决于任务因子的大量组合空间,全面详尽评估难以处理。此外,实际硬件评估缓慢且资源密集,当前做法是使用狭窄测试套件,可能遗漏关键故障模式并误判真实部署准备情况。我们提出一个主动评估框架,将策略评估视为顺序实验设计问题来应对这一挑战。我们的方法在任务因子的结构化空间上拟合概率替代模型,并自适应选择评估配置以最大化策略性能分布的信息增益,从而高效表征策略在未见条件下的行为并系统识别易失败区域。我们在3个任务、3种因子变化上进行了2331次真实世界评估,发现与典型随机测试相比,我们的方法通常能为评估者节省至少20%-40%的试验次数。
英文摘要
Generalist robot manipulation policies trained on large, diverse datasets have shown remarkable promise across a wide range of tasks. However, rigorously evaluating these policies remains a fundamental challenge. Real-world performance depends on a large combinatorial space of task factors including object poses and camera viewpoints, making full, exhaustive evaluation intractable. Additionally, real hardware evaluation is slow and resource-intensive, so current practice is to use narrow test suites that can miss critical failure modes and misrepresent true deployment readiness. We propose an active evaluation framework that addresses this challenge by treating policy evaluation as a sequential experimental design problem. Our approach fits a probabilistic surrogate model over a structured space of task factors and adaptively selects evaluation configurations to maximize information gain over the policy's performance distribution, allowing for sample-efficient characterization of policy behavior across unseen conditions and a systematic identification of failure-prone regions. We conduct 2331 real-world evaluations across 3 tasks with 3 factor variations and find that our approach typically saves the evaluator at least 20-40% of trials compared to typical random testing.
发表机构
- University of Minnesota Twin Cities(明尼苏达大学双城分校)
机构由 AI 辅助整理,请以论文原文为准。