发表机构
Seoul National University; University of Southern California; North Carolina State University; University of California Los Angeles(首尔国立大学; 南加州大学; 北卡罗来纳州立大学; 加利福尼亚大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器人学习策略部署的评估瓶颈,提出 SCAPE 框架,通过校正模拟偏差、共形预测校准不确定性,在自动驾驶等任务上提升场景预测精度与样本效率,支持细粒度部署。
AI 中文摘要
可靠的性能评估是将机器人学习策略部署到真实场景中的核心瓶颈。真实测试结果可信但成本高且难以规模化,而基于模拟的测试易于规模化但不可避免地存在现实差距(sim-to-real gap)带来的偏差。现有模拟增强方法将有限的真实场景 rollout 与大量模拟代理结合,但侧重于初始条件和部署设置的平均性能,这种群体层面的平均结果掩盖了场景特定的差异,仅能为策略可安全部署的时机和地点提供有限指导。我们提出 SCAPE,一种场景条件化的模拟增强策略评估框架,该框架利用有限的配对模拟-真实样本和大规模模拟 rollout 来预测场景条件下的真实策略性能。SCAPE 在训练预测模型前校正模拟标签中的现实差距偏差,并通过 conformal prediction(共形预测)校准预测不确定性。我们在自动驾驶和四足机器人速度跟踪任务上验证了 SCAPE。在模拟到模拟的研究中,与场景条件化神经网络基线和聚合统计基线相比,SCAPE 平均将场景层面的预测误差分别降低 4.9%/34.7%(自动驾驶)和 14.5%/27.7%(四足机器人)。我们进一步评估了部署在物理 Unitree Go2 上的速度跟踪策略,SCAPE 还提高了测试样本效率,生成更窄的校准预测区间,对分布外场景的泛化能力更强,并支持细粒度的部署策略。
英文摘要
Reliable performance evaluation is a central bottleneck for deploying robot-learning policies in real-world conditions. Real-world testing is faithful but costly and difficult to scale, whereas simulation-based testing scales easily but is inevitably biased by the sim-to-real gap. Existing simulation-augmented methods combine limited real-world rollouts with abundant simulation proxies, but focus on performance averaged over initial conditions and deployment settings. Such population-level averages obscure scenario-specific variation and provide limited guidance about when and where a policy can be safely deployed. We propose SCAPE, a scenario-conditioned simulation-augmented policy evaluation framework that predicts scenario-conditioned real-world policy performance using limited paired sim-and-real samples and large-scale simulation rollouts. SCAPE corrects sim-to-real bias in simulation labels before training the prediction model and calibrates prediction uncertainty through conformal prediction. We validate SCAPE on autonomous driving and quadruped velocity tracking. In sim-to-sim studies, SCAPE reduces scenario-level prediction error by 4.9%/34.7% (driving) and 14.5%/27.7% (quadruped) relative to scene-conditioned neural and aggregate statistical baselines on average. We further evaluate a velocity-tracking policy deployed on a physical Unitree Go2. SCAPE also improves testing sample efficiency, produces narrower calibrated prediction intervals, generalizes better to out-of-distribution scenarios, and enables fine-grained deployment strategies.
Comments22 pages