发表机构
Georgia Tech(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MOCHA提出首个基于强化学习的单网络方法,利用超网络架构生成帕累托最优策略族,并通过进化搜索高效优化机器人设计与策略组合。
AI 中文摘要
在这项工作中,我们提出了MOCHA,据我们所知,这是第一个基于强化学习的方法,能够使用单一网络在机器人的设计空间中计算一族帕累托最优策略。具体来说,MOCHA利用超网络架构学习一个网络,该网络为给定的目标和参数化的机器人设计生成优化的专用网络参数;我们将其称为多目标设计超网络(MDH)。我们在两种不同的机器人形态上展示了MDH表示复杂的设计依赖策略族的能力,每种形态具有六个设计维度,并跨越2-3个目标。此外,我们提出了一种通过对学习到的策略网络进行进化搜索来高效生成设计帕累托集的方法,为每个目标优先级生成最优的设计-策略组合。最后,我们提供了一种高效的方法来计算通用型机器人设计,该设计在整个目标集合上实现最佳的累积性能。
英文摘要
In this work, we present MOCHA, the first, to our knowledge, reinforcement learning based approach to computing a family of Pareto-optimal policies across the design space of a robot using a single network. Specifically, MOCHA leverages the hypernetwork architecture to learn a network that produces specialized network parameters optimized for a given objective and parameterized robot design; we term this a multi-objective design hypernetwork (MDH). We demonstrate the capabilities of MDHs to represent a complex family of design-dependent strategies on two distinct robot morphologies, each with six design dimensions and across 2-3 objectives. Moreover, we propose an approach for efficiently producing a Design Pareto set using evolutionary search of the learned policy network, generating the optimal design-policy combination for each objective prioritization. Lastly, we provide an efficient method for computing generalist robot designs which achieve the best cumulative performance across the entire set of objectives.