arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18827cs.LGcs.AIcs.CL

MLREF:基于大语言模型的强化学习奖励设计中高效模块复用框架

MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models

Chenglin Liu, Xun Wang, Ruishuo Chen, Zhuoran Li, Longbo Huang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对强化学习奖励设计瓶颈,提出MLREF框架,通过模块池复用奖励组件,结合三种优化机制,在17个任务上实现比强基线更优且更稳定的性能。

中文摘要 AI 辅助

奖励函数设计仍是强化学习的瓶颈。虽然大语言模型(LLMs)已支持自动奖励生成,但现有方法将奖励函数作为整体程序生成和修改,难以可靠保留和复用早期迭代中发现的有效组件,导致各迭代间性能不稳定。为解决该问题,我们提出模块级奖励演化框架(MLREF)。MLREF的核心是模块池,这是一个持久化的可复用奖励组件仓库。MLREF将模块池作为主要优化对象:池通过积累成功模块、优化表现不佳的模块、复用已验证的组件在迭代中演化;而奖励函数则由从该池抽取的模块线性组合构成。为驱动该演化,MLREF集成了三种机制:基于反思的优化、混合信用分配、带回滚的合并策略,共同提升奖励优化的有效性和鲁棒性。在17个任务上的实验表明,MLREF在运动任务中比强基线方法性能提升25.2%,在操作任务中提升6.6%,且优化动态更稳定。

英文摘要

Reward function design remains a bottleneck in reinforcement learning. While large language models (LLMs) have enabled automated reward generation, existing methods generate and revise reward functions as monolithic programs, making it difficult to reliably preserve and reuse effective components discovered in earlier iterations, leading to unstable performance across iterations. To address this, we propose Module Level Reward Evolution Framework (MLREF). At the core of MLREF is a module pool, a persistent repository of reusable reward components. MLREF treats the module pool as the primary optimization object: the pool evolves across iterations by accumulating successful modules, refining underperforming ones, and reusing proven components; while reward functions are constructed as linear combinations of modules drawn from this pool. To drive this evolution, MLREF integrates three mechanisms: reflection-based refinement, hybrid credit assignment, and a merge strategy with rollback, which together improve the effectiveness and robustness of reward optimization. Experiments on 17 tasks show that MLREF outperforms strong baselines by 25.2% in locomotion and 6.6% in manipulation, with more stable optimization dynamics.

发表机构

  • Institute for Interdisciplinary Information Sciences Tsinghua University(清华大学交叉信息研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑