arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06015cs.LGcs.AI

ProDVI:用于价值网络初始化的程序化动态先验

ProDVI: Programmatic Dynamics Priors for Value Network Initialization

Xinwei Liu, Junyuan Liang, Jianting Zhang, Wuhui Chen

AI总结:

ProDVI是利用大型语言模型生成的Python函数初始化RL智能体的框架,通过预训练价值网络编码器,提升无模型RL算法的样本效率,无需依赖数据集或模拟器。

AI中文摘要:

深度强化学习(RL)以样本效率低著称,一个促成因素是RL智能体通常从头开始初始化,迫使它们通过在线交互获取任务相关知识。现有方法通过预先收集的数据集、高保真模拟器或对相关任务的元学习来获取信息丰富的初始化,但这些前提条件可能难以获取甚至不可用。本文提出了用于价值网络初始化的程序化动态先验(ProDVI),这是一个利用大型语言模型中编码的常识和领域知识来初始化RL智能体的框架,无需依赖这些资源。具体而言,ProDVI提示代码生成型语言模型生成可执行的Python函数,这些函数编码了关于环境动态的粗略假设,然后用这些函数生成合成转移。基于这些转移,我们构建了一个辅助动态预测目标,以在演员-评论家框架中预训练价值网络的状态-动作编码器。学习到的表示在在线RL开始前提供了感知动态的归纳偏置。重要的是,生成的程序仅用于表示预训练,无需忠实地模拟目标环境。虽然生成的程序可能不准确,但它们诱导的初始化可以通过来自真实转移和奖励的在线学习进行修正。在OpenAI Gym和DeepMind Control Suite任务上的实验表明,ProDVI可有效提高无模型RL算法的样本效率。

英文摘要:

Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch, forcing them to acquire task-relevant knowledge through online interaction. Existing approaches obtain informative initializations through pre-collected datasets, high-fidelity simulators, or meta-learning over related tasks, but these prerequisites may be difficult to access or even unavailable. In this paper, we propose Programmatic Dynamics Priors for Value Network Initialization (ProDVI), a framework that leverages the commonsense and domain knowledge encoded in large language models to initialize RL agents without relying on these resources. Specifically, ProDVI prompts a code-generating language model to produce executable Python functions that encode coarse hypotheses about environment dynamics. These functions are then used to generate synthetic transitions. Based on these transitions, we construct an auxiliary dynamics prediction objective to pretrain the state-action encoder of the value network in an actor-critic framework. The learned representation provides dynamics-aware inductive biases before online RL begins. Importantly, the generated programs are used only for representation pretraining and are not required to faithfully simulate the target environment. While the generated programs may be inaccurate, their induced initialization can be corrected through online learning from real transitions and rewards. Experiments on OpenAI Gym and DeepMind Control Suite tasks show that ProDVI can effectively improve the sample efficiency of model-free RL algorithms.

↑