发表机构
University College London; Peking University; Nanyang Technological University; University of the Chinese Academy of Sciences; Beijing University of Posts and Telecommunications; University of Leeds; Fullive-AI(伦敦大学学院; 北京大学; 南洋理工大学; 中国科学院大学; 北京邮电大学; 利兹大学; Fullive-AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AquaMend提出一种基于期望损失和联合后验的探测-回滚策略,在仿真中28/32例恢复成功,平均损失较重启降低21.6%,且与DTT性能相当。
AI 中文摘要
物理变化或感知错误可能使具身智能体与任务相关的信念失效。AquaMend在探测-信念-动作图上,基于覆盖感知、物理恢复和未纠正故障的期望损失目标,比较了重新探测、回滚和支持性继续执行三种策略。联合后验引导单步策略,并辅以条件检测能力筛选。每个信念的三路最优解要求独立性、可分离性以及完全解析的探测;一般策略没有全局最优性保证。在自建仿真基准的32个配对场景中,AquaMend在28/32例中成功恢复,并将平均总损失相对于重启降低了21.6%。其与决策理论故障排除(DTT)的配对损失差异在Holm校正后无统计学显著性。相对于全候选消融,在线决策时间总体减少12.3%,但在未覆盖的后期阶段增加了3.4%。
英文摘要
Physical changes or sensing errors can invalidate embodied agents' task-relevant beliefs. AquaMend compares re-probing, rollback, and supported continuation on a probe-belief-action graph under an expected-loss objective covering sensing, physical recovery, and uncorrected failures. A joint posterior guides a one-step policy with conditional detection-power screening. The per-belief three-way optimum requires independence, separability, and fully resolving probes; the general policy has no global optimality guarantee. Across 32 paired scenarios in a self-constructed simulation benchmark, AquaMend recovers in 28/32 cases and reduces mean complete loss by 21.6% versus restart. Its paired loss difference from decision-theoretic troubleshooting (DTT) is not statistically significant after Holm correction. Against the all-candidate ablation, online decision time decreases by 12.3% overall but increases by 3.4% in the uncovered late stage.
Comments29 pages, 1 figure. Yufan Liu, Shang Luo, and Yang Liu contributed equally. Corresponding author: Bin Chong