arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于单调最优停止的预测性机器人守门

Anticipatory Robot Goalkeeping via Monotone Optimal Stopping

Hao E. Zhang, Ruize Geng, Yisen Li, Yaru Niu, Yikai Wang, Raihan Haque, Khalil Zbiss, Guanyang Luo, Hui-ping Wang, H. Eric Tseng, Ding Zhao

arXiv 2609.23976首次发表:更新:

发表机构

Carnegie Mellon University; University of Texas at Arlington; General Motors(卡内基梅隆大学; 德克萨斯大学阿灵顿分校; 通用汽车公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器人快速物理交互中需在不确定下提前行动的问题,提出单调最优停止(MOS)方法,通过强化学习训练扑救策略并学习行动与等待的回报优势来决定释放时机,显著提升守门成功率。

AI 中文摘要

在快速物理交互中,机器人往往需要在另一智能体的意图完全明确之前采取行动。预测性守门正是这一挑战的典型例证。等待虽能提供关于目标更可靠的信息,但会减少拦截的物理机会;而提前行动虽能保留可达性,却需要在不确定性下启动运动。给定一个固定的闭环扑救控制器,我们将何时启动运动的问题形式化为一个策略条件化的有限时域最优停止问题。基于此形式化,我们提出了单调最优停止(MOS),一种用于动态机器人拦截的结构化释放时机方法。四足机器人扑救策略通过强化学习训练,而MOS根据不断演化的机器人状态和目标信念决定该策略应何时被激活。MOS并非预测释放时间或仅依赖置信度,而是学习“现在行动”相对于“再等待一次观测”的回报优势。我们为这种“行动对等待”的边际推导出直接的贝尔曼递归,并仅对物理紧迫性施加单调性约束,以反映随着时间流逝拦截机会的不可逆损失。这一结构使得在动态要求高的扑救中能够早期激活,同时在后续观测改变预测目标时保留闭环自适应能力。在单一交叉条件下,MOS具有阈值释放边界且近似误差有界。大量仿真研究表明,与参数匹配的学习门控相比,MOS将平均扑救成功率从67.7%提升至74.4%,并将逆转扑救率从52.1%提升至66.5%。真实机器人实验进一步展示了在人类射门方向假动作下快速拦截和释放后方向修正的能力。

英文摘要

Robots engaged in fast physical interactions often need to act before the intent of another agent is fully known. Anticipatory goalkeeping illustrates this challenge. Waiting provides more reliable information about the target but reduces the physical opportunity for interception, whereas acting early preserves reachability but requires initiating motion under uncertainty. Given a fixed closed-loop save controller, we formulate the decision of when to initiate motion as a policy-conditional finite-horizon optimal stopping problem. Building on this formulation, we propose monotone optimal stopping (MOS), a structured release-timing method for dynamic robotic interception. The quadruped save policy is trained with reinforcement learning, while MOS determines when the policy should be activated from the evolving robot state and target belief. Rather than predicting a release time or relying on confidence alone, MOS learns the return advantage of acting now over waiting for one more observation. We derive a direct Bellman recursion for this act-versus-wait margin and impose monotonicity only with respect to physical urgency, reflecting the irreversible loss of interception opportunity as time elapses. This structure enables early activation for dynamically demanding saves while preserving closed-loop adaptation when later observations change the predicted target. Under a single-crossing condition, MOS admits a threshold release boundary with a bounded approximation error. Extensive simulation studies show that MOS improves the mean save rate from 67.7% to 74.4% over a parameter-matched learned gate and increases reversal saves from 52.1% to 66.5%. Real-robot experiments further demonstrate rapid interception and post-release direction correction under human shot-direction feints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑