发表机构
The Hong Kong Polytechnic University; The Chinese University of Hong Kong; Mohamed bin Zayed University of Artificial Intelligence; Great Bay University(香港理工大学; 香港中文大学; 穆罕默德·本·扎耶德人工智能大学; 大湾区大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对手术操作中稀疏奖励问题,提出相位与首次到达VLM反馈,通过识别最远阶段和到达时间实现信用分配,在仿真和硬件上显著提升成功率。
AI 中文摘要
稀疏的结果反馈限制了机器人从复杂操作的不成功尝试中学习的能力。失败的多阶段手术尝试可能包含值得复用的抓取、提起或转移动作。在稀疏奖励强化学习中,终止奖励将这些尝试归结为相同的结果,而标量视觉语言模型(VLM)评分既不能揭示哪些进展值得获得信用,也不能揭示进展发生的时间。我们引入了相位与首次到达反馈:对每个记录的回合进行一次VLM查询,以识别视觉上验证的最远任务阶段以及该阶段首次到达的时间,从而使学习器能够复用部分行为并定位信用。我们在SurgPhaseBench中实例化了该方法,这是一个涵盖刚性和可变形任务的相位结构化套件,并在仿真和硬件上进行了评估。在五个模拟任务中,我们的方法达到了75.2%的平均成功率,而基于对比语言-图像预训练(CLIP)并使用相同视觉输入的奖励方法仅为52.1%;当仅改变反馈表示时,这一优势仍然存在。在硬件上,相同的记录支持自主的积木拾取和滑动恢复。这些结果共同表明,轨迹级视觉监督可以保留部分进展,同时提供稀疏奖励控制所需的时间信用。
英文摘要
Sparse outcome feedback limits what robots can learn from unsuccessful attempts at complex manipulation. Failed multi-stage surgical attempts can contain grasps, lifts, or transfers worth reusing. In sparse-reward reinforcement learning, terminal rewards collapse such attempts to the same outcome, while scalar vision-language model (VLM) ratings reveal neither what progress merits credit nor when it occurred. We introduce phase-and-first-arrival feedback: one VLM query per recorded episode identifies the furthest visually verified task phase and when that phase is first reached, allowing the learner to reuse partial behavior and localize credit. We instantiate it in SurgPhaseBench, a phase-structured suite spanning rigid and deformable tasks, and evaluate it in simulation and hardware. Across five simulated tasks, our method reaches 75.2% mean success, compared with 52.1% for a reward based on Contrastive Language-Image Pre-training (CLIP) using the same visual input; the advantage persists when only the feedback representation changes. On hardware, the same record supports autonomous block picking and slip recovery. Together, these results show that trajectory-level visual supervision can preserve partial progress while providing the temporal credit needed for sparse-reward control.
Comments8 pages, 8 figures. Submitted to IEEE Robotics and Automation Letters (RA-L). Project page: https://surgphase.verloge.space