相对价值学习
Relative Value Learning
AI总结:
研究提出相对价值学习框架,通过反对称函数学习价值差异,引入成对贝尔曼算子并推导相关目标及估计器,将其与PPO集成,在Atari基准测试中取得有竞争力性能,证明相对价值估计可替代绝对评论家。
AI中文摘要:
在强化学习中,评论家通常估计绝对状态值$V(s)$,即孤立地评估特定情况有多好。然而,事实证明只有价值差异与控制相关。基于此,我们提出相对价值学习(RV)框架,它通过反对称函数$\Delta(s_i, s_j) = V(s_i) - V(s_j)$直接学习价值差异。我们引入成对贝尔曼算子并证明它是一个$\gamma$收缩,具有唯一不动点等于真实价值差异,推导出适定的单步、多步和$\lambda$回报目标,并从成对差异重建广义优势估计以获得无偏策略梯度估计器(R-GAE)。除理论结果外,我们将RV与PPO集成,在Atari基准测试(49个ALE游戏)中与标准PPO相比取得了有竞争力的性能,表明相对价值估计是绝对评论家的有效替代方案。
英文摘要:
In reinforcement learning, critics typically estimate absolute state values $V(s)$, estimating how good a particular situation is in isolation. However, it turns out that only differences in value are relevant for control. Motivated by this, we propose Relative Value Learning (RV), a framework that learns value differences directly via an antisymmetric function $Δ(s_i, s_j) = V(s_i) - V(s_j)$. We introduce a pairwise Bellman operator and prove it is a $γ$-contraction with a unique fixed point equal to the true value differences, derive well-posed $1$-step, $n$-step and $λ$-return targets and reconstruct generalized advantage estimation from pairwise differences to obtain an unbiased policy-gradient estimator (R-GAE). Beyond theoretical results, we integrate RV with PPO and achieve competitive performance on the Atari benchmark (49 ALE games) compared to standard PPO, indicating that relative value estimation is an effective alternative to absolute critics.