arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21572cs.RO

为仿真到真实性能证书下注

Betting for Sim-to-Real Performance Certificates

Yujia Chen, Bowen Weng

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出仿真到真实下注证书框架,通过算法关联模拟器库与赌注,证明证书任意时刻有效,实验显示其能大幅缩窄机器人系统的性能证书区间。

中文摘要 AI 辅助

考虑机器人系统的典型测试:观测关于某一关注方面的一系列结果(碰撞或无碰撞、跟踪误差、完成时间),并报告均值(碰撞风险、平均误差、平均完成时间),更重要的是,报告一个以规定置信度保证包含该均值的区间,即性能证书。由于真实世界试验成本高昂,样本量通常很小,证书往往较为宽松。现在考虑相同的流程,只是在每个真实结果公布前,操作者会“偷看”大量模拟结果,并对真实结果将落在何处下注。随着真实结果确定赌注,操作者的财富会增减,对模拟器的“信任”也会在组合中发生变化。本文将这一思路发展为仿真到真实下注证书框架,有三个贡献:(i)一种算法,将可扩展的模拟器库与有效赌注关联,并将累积的赌注财富与证书关联;(ii)证明返回的证书是任意时刻有效的,使用任何模拟器库都能以规定概率覆盖真实均值;(iii)保证的财富遗憾边界为所提算法和模拟器库设计提供配置原则,以生成更紧凑的证书。在合成分布和真实机器人测试上的实验,涵盖重放的标准化测试结果和在线运行时评估,表明所提方法相较于经典和最先进基线,能将证书缩小$51.6\rm\textit{\textpm}16$%,在样本量极少的情况($\text{样本数}\text{\textless}\text{=}30$)下则能缩小$32.26\rm\textit{\textpm}8$%。

英文摘要

Consider a typical test of a robot system: one observes a sequence of outcomes concerning some aspect of interest (crash or no crash, tracking error, time to completion), and reports a mean (crash risk, average error, mean time to completion) and, more importantly, an interval guaranteed to contain that mean at a prescribed confidence, referred to as a performance certificate. Given expensive real-world trials, the sample size is therefore small, and the certificate is often loose. Now consider the same procedure, except that before each real outcome is revealed, the operator ``peeks'' at a large bank of simulated results, and places a bet on where the real outcome will land. As the real outcomes settle the bets, the operator gains or loses wealth. One's ``trust'' over simulators also shifts within the portfolio. This paper develops that idea into a sim-to-real betting certificate framework with three contributions: (i) An algorithm that links a scalable bank of simulators to effective bets, and the accumulated betting wealth to the certificate. (ii) A proof that the returned certificate is anytime valid, covering the true mean with the prescribed probability, using any simulator bank. (iii) The guaranteed wealth-regret bounds yield configuration principles for the proposed algorithm and simulator bank design to deliver tight certificates. Experiments across synthetic distributions and real-world robot tests, covering both replayed standardized testing outcomes and online runtime evaluation, show the proposed method narrows the certificate by $51.6\%\pm16\%$ against classic and state-of-the-art baselines, and by $32.26\%\pm8\%$ in the extremely limited-sample regime ($\leq30$ samples).

发表机构

  • Iowa State University(爱荷华州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑