arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越孤立:解锁强化学习组件协同以实现样本高效的连续控制

Beyond Isolation: Unlocking Reinforcement Learning Component Synergy for Sample-Efficient Continuous Control

Qi Zhao, Guozheng Ma, Yilun Kong, Lu Li, Haoyu Wang, Zilin Wang, Tiantian Zhang, Yuxing Wang, Jian Sha, Yongzhe Chang, Xueqian Wang, Dacheng Tao

arXiv 2608.07086首次发表:更新:

发表机构

Tsinghua University; Nanyang Technological University; Mila - Quebec Artificial Intelligence Institute; University of Oxford(清华大学; 南洋理工大学; 米拉-魁北克人工智能研究所; 牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对强化学习组件孤立问题,提出协调模型表示等三个维度的ROSER框架,在连续控制基准上性能优于基线及单纯堆叠方法,为样本高效智能体开发提供了整体设计思路。

AI 中文摘要

强化学习(RL)系统因固有属性而比其他机器学习范式复杂得多,导致RL系统设计需同时考虑许多紧密耦合的因素。尽管单个算法组件已取得进展,但它们的功能相互依赖性仍未被探索:它们表现出相互协同还是反生产性干扰?为弥合这一差距,我们进行了系统研究,发现不同组件的功效具有显著的任务依赖性,而单纯堆叠最先进技术不一定能带来性能提升;相反,它常引发复合非平稳性等突发挑战。基于这些发现,我们提炼出一套用于这些组件原则性协调的可操作见解。在这些见解指导下,我们提出ROSER,这是一个协调三个关键维度的RL框架:基于模型的表示、优化稳定性和经验回放。在多种连续控制基准测试中,ROSER始终优于普通基线,且比单纯堆叠方法实现了17.60%的性能提升。我们的研究结果强调了RL系统设计中整体视角的必要性,并为开发样本高效智能体铺平了道路。

英文摘要

Reinforcement learning systems are significantly more complex than other machine learning paradigms due to inherent properties, causing RL system design to jointly account for many tightly coupled factors. Despite advances in individual algorithmic components, their functional interdependencies remain underexplored: do they exhibit mutual synergy or counterproductive interference? To bridge this gap, we conduct a systematic investigation and find that the efficacy of different components exhibits significant task-dependency, and naively stacking state-of-the-art techniques does not necessarily yield performance gains; instead, it often triggers emergent challenges, such as compounded non-stationarity. Building upon these findings, we distill a suite of actionable insights into the principled coordination of these components. Guided by these insights, we propose ROSER, an RL framework that coordinates three critical dimensions: Model-based Representation, Optimization Stability, and Experience Replay. Across diverse continuous-control benchmarks, ROSER consistently outperforms vanilla baselines and achieves 17.60% gains over naive stack. Our findings underscore the necessity of a holistic perspective in RL system design and paves the way for developing sample-efficient agents.

Comments27 pages including appendix, 10 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑