发表机构
Carnegie Mellon University; University of Texas at Arlington(卡内基梅隆大学; 德克萨斯大学阿灵顿分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出动力学诱导承诺映射(DIC-Map)分析人形-四足罚球系统中的身体-策略耦合,发现承诺出现在接触前0.29秒,替换估计器可显著提升扑救率至0.472。
AI 中文摘要
机器人游戏中的学习不仅受到策略信息的约束,还受到身体仍能执行的动作的约束。我们在一个分层的人形-四足罚球系统中研究这种耦合,其中游戏层面的自对弈策略命令固定的足球全身控制器(S-WBCs)。人形射门技能从自收集的动作捕捉数据初始化,而四足扑救技能通过强化学习获得。我们引入了动力学诱导的承诺映射(DIC-Map),这是一种基于身体的接地分析,用于估计继续能力,识别终端替代方案的第一个持续丧失,并测试剩余交互是否允许简化的零和博弈。对于对称的终端替代方案,简化博弈得出了由响应者延迟值决定的最优策略集中度的闭式界。我们进一步表明,当响应者通过估计器行动时,相等的响应值消除了直接的终端分配梯度,并留下一个由估计器介导的一阶学习通道。实验定位承诺在接触前约0.29秒,仅改变球速就会改变延迟覆盖范围。在四种响应者策略中,替换估计器将扑救率从0.240提高到0.472,而通过等待获得的相当的可读性增益仅将其提高到0.246。由于可用的覆盖项是观察性代理,因此使用后验分析进行均衡比较。项目网站:此 https URL
英文摘要
Learning in robotic games is constrained not only by strategic information but also by what the body can still execute. We study this coupling in a hierarchical humanoid-quadruped penalty system in which game-level self-play policies command fixed soccer whole-body controllers (S-WBCs). The humanoid shooting skill is initialized from self-collected motion-capture data, whereas the quadruped saving skill is learned by reinforcement learning. We introduce dynamics-induced commitment mapping (DIC-Map), a body-grounded analysis that estimates continuation capability, identifies the first persistent loss of a terminal alternative, and tests whether the remaining interaction admits a reduced zero-sum game. For symmetric terminal alternatives, the reduced game yields a closed-form bound on optimal strategy concentration determined by the responder's value of deferring. We further show that, when the responder acts through an estimator, equal response values eliminate the direct terminal-allocation gradient and leave an estimator-mediated first-order learning channel. Experiments locate commitment about 0.29 s before contact, and changing only ball speed shifts deferral coverage. Across four responder policies, replacing the estimator raises save rate from 0.240 to 0.472, whereas a comparable gain in read accuracy obtained by waiting raises it only to 0.246. Posterior analysis is used for the equilibrium comparison because the available coverage terms are observational proxies. Project website: https://chris-ruizegeng.github.io/penaltykick/