arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于可扩展强化学习的简单演员与深度评论家

Simple Actors and Deep Critics for Scalable Reinforcement Learning

Guhyeon Kang, Jaehwi Lee, Minhae Kwon

arXiv 2608.26659首次发表:更新:

发表机构

Sungkyunkwan University; Soongsil University(成均馆大学; 崇实大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出LAC算法,将容量分配给深度评论家而非简单演员,解决离线RL中加深评论家的三种失效模式,在OGBench上实现与强基线相当性能且推理延迟最高降4倍。

AI 中文摘要

离线强化学习(RL)的近期进展得益于扩散、流匹配策略等表达性生成演员,这些演员能捕捉离线数据集中的多模态行为。但这类演员每步动作需多次去噪或积分,部署时每步决策都会产生大量开销。本文重新探讨离线演员-评论家方法应如何分配容量:由于评论家仅在训练时使用,部署时会被丢弃,而演员需在每步决策运行,因此将容量分配给评论家而非演员更利于推理效率。不过,离线RL中扩大MLP评论家规模会引发多种明显不稳定性,实际应用中评论家多为浅层。我们确定了离线RL中加深评论家时出现的三种不同失效模式:优化失效、自举噪声放大、价值范围漂移,并分别用残差MLP主干、n步自举目标、分类交叉熵损失解决。将这些成分与轻量确定性演员结合,我们提出LAC(Light Actor, deep Critic)。在OGBench上,LAC与最强的扩散、流匹配基线性能相当,推理延迟最高降低4倍,与无蒸馏的单步蒸馏策略相当;其评论家配方可跨演员参数化迁移。

英文摘要

Recent progress in offline reinforcement learning (RL) has been driven by expressive generative actors such as diffusion and flow-matching policies, which capture multimodal behavior in offline datasets. However, these actors require multiple denoising or integration steps per action and thus incur substantial overhead at every decision in deployment. In this work, we revisit where capacity should be invested in an offline actor--critic method. Since the critic is used only during training and is discarded at deployment while the actor runs at every decision step, allocating capacity to the critic rather than the actor is more favorable for inference-time efficiency. However, scaling MLP critics in offline RL is known to introduce several distinct instabilities that have, in practice, kept critics shallow. We identify three distinct failure modes that arise when critics are deepened in offline RL---optimization, bootstrap-noise amplification, and value-range drift---and address each with a corresponding ingredient: a residual MLP backbone, n-step bootstrap targets, and a categorical cross-entropy loss. Combining these ingredients with a lightweight deterministic actor, we propose LAC (Light Actor, deep Critic). On OGBench, LAC matches the strongest diffusion- and flow-matching baselines while achieving up to 4x lower inference latency, comparable to one-step distilled policies without distillation. Its critic recipe also transfers across actor parametrizations.

CommentsAccepted at CIKM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑