arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

方差厌恶的 $n$ 步离线强化学习用于稀疏长时域环境

Variance-Averse $n$-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments

Guhyeon Kang, Minhae Kwon

arXiv 2610.07899首次发表:更新:

发表机构

Sungkyunkwan University(成均馆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对生成式离线强化学习在异构数据中高方差动作不可靠的问题,提出VAN-Flow框架,结合分类分布评论家、方差厌恶期望算子与流匹配生成行动者,在40余个任务上优于强基线,尤其提升长时域场景性能。

AI 中文摘要

生成式行动者正在通过建模复杂动作分布的强表达力策略类别来变革离线强化学习(RL)。然而,这种表达力也暴露了异构数据集中的一个关键挑战:生成式策略可能复现不可靠的动作模式,这些模式的回报分布具有高方差,偶尔因偶然因素获得高回报但缺乏一致性。因此,仅最大化期望 $Q$ 值不足以识别可靠动作。我们提出 VAN-Flow(方差厌恶的 $n$ 步流),一个在生成式离线 RL 中促进可靠动作的框架。VAN-Flow 结合了(i)一个分类分布评论家,(ii)一个方差厌恶的期望算子,该算子平滑地重新加权原子概率以偏向具有高回报和低离散度的动作,以及(iii)一个通过拒绝采样引导的流匹配生成式行动者。与 CVaR 或均值-方差目标不同,该算子在不进行硬截断或添加辅助惩罚项的情况下,在分类回报分布上重新分配概率质量。在来自 D4RL 和 OGBench 的超过 40 个任务中,VAN-Flow 持续优于强基线,在长时域和高方差场景中改进最大,此时可靠动作选择变得至关重要。

英文摘要

Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency. Consequently, maximizing the expected $Q$-value alone is insufficient for identifying reliable actions. We propose VAN-Flow (Variance-Averse $n$-step Flow), a framework that promotes reliable actions in generative offline RL. VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling. Unlike CVaR or mean-variance objectives, the operator redistributes probability mass over the categorical return distribution without hard truncation or auxiliary penalty terms. Across more than 40 tasks from D4RL and OGBench, VAN-Flow consistently outperforms strong baselines, with the largest gains in long-horizon and high-variance regimes where reliable action selection becomes critical.

CommentsAccepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑