arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MEMENTO:内存引导的基于代码的策略进化的模因算法

MEMENTO: Memory-Guided Memetic Code-as-Policy Evolution

Alkis Sygkounas, Victor Aregbede, Amy Loutfi, Andreas Persson

arXiv 2607.22832首次发表:更新:

发表机构

Center for Applied Autonomous Sensor Systems (AASS), Örebro University; Örebro University(应用自主传感器系统中心(AASS),于默奥大学; 于默奥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究长期具身任务中基于代码的策略进化问题,提出内存引导单精英模因框架MEMENTO,通过进化展开评估器及利用反馈指标改进策略,在相关领域实验中性能优于基线,还证明了模拟到现实转移的可行性。

AI 中文摘要

长期的具身任务需要策略在任务成功之前执行许多相关动作。将策略表示为可执行控制程序(基于代码的策略)能够在展开评估后检查和修改其决策逻辑。通过展开性能比较修订后的程序,将策略改进框架化为执行引导的程序搜索。由大语言模型驱动的进化方法通过生成变体和选择高性能候选者为这种搜索提供了一种自然机制。然而,现有方法主要在独立生成的变体中进行选择,缺乏顺序局部改进阶段。我们引入了MEMENTO,一种用于基于代码的策略进化的内存引导单精英模因框架。MEMENTO首先进化一个展开评估器,将策略展开映射到标量适应度和结构化反馈指标。适应度选择被接受的候选者和下一个精英,而反馈指标则限制由内存引导的爬山、宏变异和交叉生成的策略提议。我们在两个长期具身领域评估了MEMENTO:Robosuite Franka汉诺塔操作和AI2-THOR家庭交互。在任务成功率以及对保留的Robosuite对象配置和未见的AI2-THOR场景的泛化方面,MEMENTO优于作为基于代码的策略进化基线的Eureka和REvolve。消融实验表明,零样本生成和未进化的评估器无法解决任何一个领域的问题,并且去除策略搜索分支会降低性能。最后,我们在物理Franka机器人上部署了最佳进化的Robosuite策略,证明了进化后的基于代码的策略从模拟到现实转移的可行性。代码、提示和视频可在:此https URL获取。

英文摘要

Long-horizon embodied tasks require policies that execute many dependent actions before task success can be observed. Representing policies as executable control pro- grams (code-as-policy) enables their decision logic to be inspected and revised after rollout evaluation. Revised programs can then be executed and compared by rollout performance, framing policy improvement as execution-guided program search. Evo- lutionary methods driven by large language models (LLMs) provide a natural mecha- nism for this search by generating variants and selecting high-performing candidates. However, existing approaches primarily select among independently generated vari- ants and lack a sequential local improvement phase. We introduce MEMENTO, a memory-guided single-elite memetic framework for code-as-policy evolution. ME- MENTO first evolves a rollout evaluator that maps policy rollouts to scalar fitness and structured feedback metrics. Fitness selects accepted candidates and the next elite, while feedback metrics condition policy proposals generated by memory-guided hill-climbing, macro-mutation, and crossover. We evaluate MEMENTO on two long- horizon embodied domains: Robosuite Franka Tower-of-Hanoi manipulation and AI2- THOR household interaction. MEMENTO outperforms Eureka and REvolve, adapted as code-as-policy evolutionary baselines, in task success and generalization to held- out Robosuite object configurations and unseen AI2-THOR scenes. Ablations show that zero-shot generation and unevolved evaluators fail to solve either domain, and that removing policy-search branches reduces performance. Finally, we deploy the best-evolved Robosuite policy on a physical Franka robot, demonstrating the feasibil- ity of sim-to-real transfer of the evolved code-as-policy. Code, prompts, and videos are available at: https://github.com/sygkounas/MEMENTO.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑