超越最终任务成功率:如何审计机器人中的视觉经验检索
Beyond End-Task Success: How to Audit Visual Experience Retrieval in Robotics
浏览论文内容
中文总结 AI 辅助
本文提出一种审计机器人视觉经验检索的方法,通过穷举执行所有经验来区分选择规则与经验库质量,发现视觉规则易集中于单一经验且排序能力有限,并建议报告经验分布与事后最佳经验成功率。
中文摘要 AI 辅助
存储过往经验的机器人在新场景中必须选择复用哪一条经验。大多数系统通过视觉相似性进行选择,而大多数评估仅报告所选经验的成功率。该数字无法说明选择是否良好:一个规则可能因反复使用一条广泛可迁移的经验而得分很高,也可能因其偏好的经验较弱而得分很低。由于机器人越来越多地通过复用而非重新训练来适应,描述经验库而非规则的分数会误导领域后续的构建方向。我们提出一种审计方法:在两个操作任务、三种复用机制以及经验库规模为$K=3$、$10$和$50$的情况下,在每一个查询场景中执行每一条存储的经验。由于所有备选方案的结果已知,分数可追溯至逐场景选择或经验库质量。被审计的规则在五种视觉嵌入(从原始像素到CLIP)中依据最近邻距离进行选择。(1) 一条用事后视角选定的固定经验,可捕获随机选择与理想选择之间差距的30-58%;逐场景选择则在成功率上竞争剩余的0.07-0.15。(2) 当$K\ge10$时,视觉规则对某一条经验的集中程度是理想选择的1.5-3倍,其分数随后取决于该经验的质量。(3) 无论何处,若某规则显著不同于保持其选择率但将场景随机配对的洗牌版本,则该规则更差,且对所有学习到的图像策略均如此。(4) 视觉距离能很好地预测给定配对是否成功(AUROC最高达0.96),但在$K=50$时,对五种嵌入中的四种,其在单场景内对候选的排序不优于随机(AUROC为0.45-0.52)。穷举执行通常不可行,因此审计简化为任何研究都能提供的两个廉价报告:所选经验的分布,以及事后最佳单条经验的成功率。
英文摘要
Robots that store past experiences must select which one to reuse in a new scene. Most systems select by visual similarity, and most evaluations report only the success of the selected experience. That number does not show whether the selection was good: a rule can score well by repeatedly using one broadly transferable experience, or poorly because its preferred experience is weak. Since robots increasingly adapt by reuse rather than retraining, a score that describes the library rather than the rule misleads what the field builds next. We contribute an audit methodology: execute every stored experience in every query scene, over two manipulation tasks, three reuse mechanisms, and libraries of $K=3$, $10$, and $50$. Because every alternative's outcome is known, a score can be traced to per-scene selection or to library quality. The audited rules select by nearest-neighbor distance in five visual embeddings, from raw pixels to CLIP. (1) One fixed experience, chosen with hindsight, captures 30-58% of the gap between random selection and an oracle; per-scene selection competes for the remaining 0.07-0.15 in success rate. (2) At $K\ge10$, visual rules concentrate on one experience 1.5-3 times more than the oracle does, and their scores then follow that experience's quality. (3) Wherever a rule differs significantly from a shuffle that keeps its selection rates but pairs them with scenes at random, the rule is worse, for every learned image policy. (4) Visual distance predicts well whether a given pair will succeed (AUROC up to 0.96), yet ranks the candidates within one scene no better than chance for four of five embeddings at $K=50$ (AUROC 0.45-0.52). Exhaustive execution is usually infeasible, so the audit reduces to two cheap reports any study can give: the distribution of selected experiences, and the success of the best single experience in hindsight.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Centific
机构由 AI 辅助整理,请以论文原文为准。