发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
InterEvolve提出测试时进化框架,通过对象感知前向-后向模型和LLM奖励程序进化,在无需重训练下释放人形机器人控制器既有技能,解决新移动操作任务。
AI 中文摘要
我们研究人形机器人移动操作的测试时进化:通过重新利用控制器从未训练过的现有技能、从自身尝试中改进并保留所学内容,在不进行重新训练的情况下解决新任务。我们的关键见解是,一个广泛的控制器已经具备新任务所需的大部分能力,而这种能力可以通过规划与控制之间的接口被访问,该接口既足够表达接触丰富、多阶段的交互,又足够可执行和可测量,从而使执行反馈能够指导基于经验的规划。InterEvolve通过两个组件实现这一接口。首先,我们开发了一个对象感知的前向-后向(FB)行为基础模型,其在冻结的身体先验上的对象残差将关于身体或对象的新奖励在测试时转化为移动操作行为。其次,我们将任务指定为奖励程序:带有完成条件和可调常数的分阶段奖励。一个大型语言模型(LLM)代理根据执行反馈和已验证程序的技能库,在上下文中修改程序结构,而数值优化器调整其常数。随着每个候选在并行模拟场景中验证,程序探索新的方式来诱导、重新利用和组合控制器现有的运动能力以完成当前任务,从而在迭代中改进。实验表明,人工设计的奖励未能充分利用FB模型的移动操作能力,而InterEvolve进化的程序释放了这些能力,有时通过新颖的策略实现。它进一步在模拟中为多样任务、复杂场景和长时程组合生成行为,进化的技能可在物理Unitree G1上从以自我为中心的机载感知自主运行。
英文摘要
We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.
CommentsProject page: https://sirui-xu.github.io/InterEvolve