arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25297cs.LG

修正优先回放中的组内自选择偏差

Correcting Within-Group Self-Selection Bias in Prioritized Replay

Oscar Miró López-Feliu, Herke van Hoof

首次发表
浏览论文内容

中文总结 AI 辅助

针对优先回放中组内自选择偏差,提出兄弟感知回放方法,通过兄弟采样、平均或模型采样修正结果分布,提升随机环境下的学习效率。

中文摘要 AI 辅助

优先经验回放(PER)通过回放高优先级转换(通常依据绝对时序差分误差)来提高样本效率。在随机环境中,PER 可能会扭曲从具有相同状态-动作对的转换中回放的实际结果的分布。我们称之为组内自选择。我们量化了由此导致的组内结果频率和平均贝尔曼目标的变化。我们将 PER 分解为组间分配和条件兄弟选择,并推导出保留当前组级优先级质量的固定缓冲区修正:SAMPLE 通过 PER 选择一个组,并对均匀采样的兄弟样本进行训练;AVG 对兄弟贝尔曼目标取平均;MODEL 从经验全结果模型中采样。在具有罕见高幅度结果的精确状态-动作环境中,兄弟感知回放比 PER 提高了学习效率,尽管匹配的参数扫描表明调参可以缩小一些差距。在 MinAtar 中,使用近似 VQ-VAE 组的 SAMPLE 在五个游戏中的四个中缓解了均值保持奖励尾部下的性能退化。因此,兄弟感知回放保留了对高优先级状态-动作区域的关注,同时恢复了其经验结果频率。

英文摘要

Prioritized experience replay (PER) improves sample efficiency by replaying high-priority transitions, usually according to absolute temporal-difference error. In stochastic environments, PER can distort the distribution of realized outcomes replayed from transitions with the same state-action pair. We call this within-group self-selection. We quantify the resulting changes in within-group outcome frequencies and mean Bellman targets. We decompose PER into between-group allocation and conditional sibling selection, and derive fixed-buffer corrections that preserve current group-level priority mass: SAMPLE selects a group through PER and trains on a uniformly sampled sibling; AVG averages sibling Bellman targets; and MODEL samples from an empirical full-outcome model. In exact state-action environments with rare high-magnitude outcomes, sibling-aware replay improves learning efficiency over PER, although matched parameter sweeps show that tuning can narrow some gaps. In MinAtar, approximate VQ-VAE groups with SAMPLE mitigate degradation under mean-preserving reward tails in four of five games. Sibling-aware replay thus retains the focus on high-priority state-action regions while recovering their empirical outcome frequencies.

发表机构

  • University of Amsterdam(阿姆斯特丹大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑