arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越任务奖励:一种用于评估具身依赖能力的控制器限制协议

Beyond Task Reward: A Controller-Restriction Protocol for Evaluating Embodiment-Dependent Competence

Siyuan Zhang

arXiv 2610.07629首次发表:更新:

发表机构

University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种控制器限制协议,通过限制控制器评估机器人形态的具身能力,弥补任务奖励无法区分控制器依赖度的缺陷,并在EvoGym任务上验证其可靠性与有效性。

AI 中文摘要

协同设计方法联合优化机器人的身体和控制器,并通过一个数字——完全优化配对的任务奖励——来评判结果。该数字无法区分那些能力对控制器依赖程度差异很大的形态。我们转而通过限制其控制器来评估一种形态,记录其在明确声明、低复杂度的控制器族、环境、任务和搜索预算下保留的任务能力。在三个EvoGym运动任务上,任务奖励仅解释了该量的$39\%$、$33\%$和$10\%$的方差,并且在按运行组分组的交叉验证下,几何描述符无法预测该量。该测量在优化器重启间是可靠的(ICC$(2,k) = 0.956$--$0.986$),但依赖于所声明的控制器族:按执行器索引而非位置进行驱动相位化,对相同形态的Spearman排序为$0.50$--$0.63$,并逆转了奖励匹配的对。作为第二个搜索目标,该轴在$5$次配对运行中的$3$次中,在匹配任务奖励下提高了能力,未达到预先注册的$4$次门槛。相反,在普通的仅奖励搜索之后使用它,用于在其最高任务奖励带内进行选择,它在所有$8$次提供选择的运行中选择了不同的身体,代价至多为$0.10$个奖励单位,并且在$8$次中的$6$次中,该身体在保留的控制器族下也获得了更高的分数。因此,受限控制能力是协同设计形态的一种可报告属性,只能通过定义它的控制器族来解释。

英文摘要

Co-design methods optimize a robot's body and controller jointly and judge the result by one number, the task reward of the fully optimized pair. That number cannot separate morphologies whose competence depends on the controller to very different degrees. We evaluate a morphology by restricting its controller instead, recording the task competence it retains under an explicitly declared, low-complexity controller family, environment, task and search budget. On three EvoGym locomotion tasks, task reward explains only $39\%$, $33\%$ and $10\%$ of the variance in this quantity, and geometric descriptors do not predict it under run-grouped cross-validation. The measurement is reliable across optimizer restarts (ICC$(2,k) = 0.956$--$0.986$) but depends on the declared family: phasing the drive by actuator index instead of position ranks the same morphologies at Spearman $0.50$--$0.63$ and reverses reward-matched pairs. As a second search objective the axis improved competence at matched task reward in $3$ of $5$ paired runs, short of a pre-registered bar of $4$. Used after an ordinary reward-only search instead, to choose within its top task-reward band, it selected a different body in all $8$ runs offering a choice, at a cost of at most $0.10$ reward units, and in $6$ of $8$ that body also scored higher under a held-out family. Restricted-control competence is therefore a reportable property of a co-designed morphology, interpretable only with the controller family that defines it.

Comments8 pages, 6 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑