arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38733cs.AIcs.LG

代码控制:合成参数化反应式控制器

Code to Control: Synthesizing Parameterized Reactive Controllers

Zergham Ahmed, Joshua B. Tenenbaum, Chris Bates, Samuel J. Gershman

首次发表
浏览论文内容

中文总结 AI 辅助

提出Code to Control方法,通过LLM合成控制器结构、无导数搜索调参,生成无需推理或规划的Python策略,在Atari、Flappy Bird和MuJoCo任务中优于规划式程序合成,媲美深度强化学习且更高效。

中文摘要 AI 辅助

近期基于LLM的控制方法要么调用语言模型来选择动作,要么合成需要在每一步决策时进行规划的世界模型,这引入了可能限制实时使用的延迟。我们提出了Code to Control,一种合成Python控制器的方法,这些控制器直接作为策略执行。Code to Control将程序结构与参数分离。LLM合成控制器结构,而无导数搜索则利用环境反馈来拟合其参数以进行连续控制。一旦学习完成,所得的控制器在决策时既不需要LLM推理也不需要规划,从而实现了实时游戏,并且在我们的计时协议下,动作选择比PPO策略更快。在一系列Atari游戏、Flappy Bird和MuJoCo任务中,Code to Control优于基于规划的程序合成方法,与深度强化学习相比保持竞争力,同时使用更少的环境交互,能够适应环境动态的显著变化,并扩展到复杂的运动任务。

英文摘要

Recent LLM-based approaches to control either invoke a language model to select actions or synthesize world models that require planning at every decision, introducing latency that can limit real-time use. We introduce Code to Control, an approach that synthesizes Python controllers which execute directly as policies. Code to Control separates program structure from parameters. An LLM synthesizes the controller structure, while derivative-free search fits its parameters for continuous control using feedback from the environment. Once learned, the resulting controllers require neither LLM inference nor planning at decision time, enabling real-time gameplay and, under our timing protocol, faster action selection than a PPO policy. Across a suite of Atari games, Flappy Bird, and MuJoCo tasks, Code to Control outperforms planning-based program synthesis methods, remains competitive with deep reinforcement learning while using fewer environment interactions, transfers across substantial changes in environment dynamics, and scales to complex locomotion tasks.

发表机构

  • Harvard University(哈佛大学)
  • Massachusetts Institute of Technology(麻省理工学院)
  • Florida Institute for Human and Machine Cognition(佛罗里达人类与机器认知研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑