arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一次三步:对比强化学习中从动作序列学习表示

Three Steps at a Time: Learning Representations from Action Sequences in Contrastive RL

Michal Korniak, Kamil Dybek, Benjamin Eysenbach, Marco Bagatella, Michał Bortkiewicz

arXiv 2608.30640首次发表:更新:

发表机构

ETH Zurich; University of Warsaw; Princeton University; MPI IS Tubingen; Warsaw University of Technology; IDEAS Research Institute(苏黎世联邦理工学院; 华沙大学; 普林斯顿大学; 图宾根马克斯·普朗克智能系统研究所; 华沙理工大学; IDEAS研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对对比强化学习(CRL)将动作建模尺度从单步扩展至动作块,在18个与11个环境中分别实现31.7%、93.1%的性能提升,揭示动作块可改进CRL的评论者表示以提升算法效果。

AI 中文摘要

尽管自监督强化学习方法通过学习状态与动作的表示已取得优异成果,但一个关键未决问题是应建模的动作时间尺度。不同于依赖单步动作的标准设定,我们将典型自监督方法对比强化学习(CRL)扩展为在动作块上运行,发现在既定离线与在线基准中均取得广泛且显著的性能提升:在18个环境中提升31.7%,在11个环境中提升93.1%。通常,动作块化带来的增益可通过其建模非马尔可夫、时间扩展策略及传播无偏多步回报的能力解释,但有趣的是,这些论点仅部分适用于CRL。我们的实证研究表明,在CRL语境下,动作块比单个动作携带更多关于目标的信息,可显著改进评论者的表示,使算法效果大幅提升。

英文摘要

While self-supervised approaches to reinforcement learning have achieved strong results by learning representations of states and actions, a key open question is the time scale over which actions should be modeled. Departing from the standard formulation relying on single-step actions, we extend contrastive reinforcement learning (CRL), a prototypical self-supervised method, to operate over action chunks, and find that this results in large, pervasive gains across established offline and online benchmarks: +31.7% and +93.1% across 18 and 11 environments respectively. While action-chunking-driven gains are generally explained through the ability to model non-Markovian, temporally extended policies, and to propagate unbiased multi-step returns, interestingly, we find that these arguments only partially apply to CRL. Our empirical studies suggest that, in the context of CRL, an action chunk carries more information about the goal than a single action, measurably improving the critic's representations, and rendering the algorithm significantly more effective.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑