先审计再执行:在一次性操纵的主动识别中定位信念失败
Audit Before You Commit: Locating Belief Failures in Active Identification for One-Shot Manipulation
浏览论文内容
中文总结 AI 辅助
本文提出在一次性操纵前,应分别审计信念覆盖与失败模型两个条件,通过离线审计定位观测模型误差(如2.1毫米敲击边界误差),修正后显著降低失败率,并在物理机械臂上验证了该方法。
中文摘要 AI 辅助
一个机器人在执行一次不可逆动作(例如插入销钉前轻敲表面)之前会进行几次探测,它必须决定何时证据足以承诺执行。我们认为,这一决定依赖于现有方法未区分的两个条件:信念必须仍然覆盖决定动作的坐标中的真实状态,并且对动作进行评分的失败模型必须追踪实际发生的失败。我们分别离线并在有真实标签的情况下,在一个已部署的“探测-承诺”流程上审计这两个条件:该流程包括一个粒子信念、一个场景失败评分和一次承诺。在模拟插入中,更多的敲击使信念更集中,而真实状态在16.9%的情节中离开了其支撑集,失败评分变得乐观了0.31。保形校准恢复了覆盖率,但没有恢复决策:自信但错误的实例仍然通过置信门。审计的特征反而指向观测模型,其中一次手动扫描发现敲击边界存在2.1毫米的误差。修正这一个数字,在未触及的实例上将失败率从0.354降至0.112,并且无需重新拟合即可迁移到第二个引擎,而在第三个引擎中,同样的审计则提示执行模型不匹配。在三个引擎的七个任务族中,在固定执行器上进行少量探测减少了未命中或失败。在一台通过触觉将工具插入刚性凹槽的物理机械臂上,增益和审计的两个条件得以重现,并且在注入模型误差的情况下重放记录的敲击,在真实数据上显示了审计的特征。附加材料可在该https URL获取。
英文摘要
A robot that probes a few times before one irreversible action, such as tapping a surface before inserting a peg, must decide when the evidence is enough to commit. We argue that this decision rests on two conditions that existing methods do not separate: the belief must still cover the truth in the coordinate that decides the action, and the failure model that scores actions must track realized failure. We audit both conditions separately, offline and with ground truth, on a deployed probe-then-commit pipeline: a particle belief, a scenario failure score, and one commit. On simulated insertion, more taps sharpen the belief while the truth leaves its support on 16.9% of episodes and the failure score turns optimistic by 0.31. Conformal calibration restores coverage but not the decision: confidently wrong instances still pass a confidence gate. The audit's signatures instead point at the observation model, where a hand scan finds a 2.1 mm error in the tap boundary. Correcting that one number cuts failure from 0.354 to 0.112 on untouched instances and transfers unrefitted to a second engine, while in a third engine the same audit suggests an execution-model mismatch instead. Across seven task families in three engines, a few probes at a fixed executor reduce miss or failure. On a physical arm inserting a tool into a rigid pocket by touch, the gain and the audit's two conditions reproduce, and replaying the recorded taps under an injected model error shows the audit's signature on real data. Additional materials are available at https://sites.google.com/view/auditbeforeyoucommit.