arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26918cs.LGcs.AI

多目标强化学习中事后重标记导致的偏好覆盖坍缩

On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning

  • Mila (Quebec AI Institute)(米拉(魁北克人工智能研究所))
  • Université Laval(拉瓦尔大学)
  • CIFAR AI Chair(CIFAR人工智能教席)

机构由 AI 辅助整理,请以论文原文为准。

Baptiste Bonin, Caro Strickland, Audrey Durand

AI总结:

发现事后重标记在多目标强化学习中导致偏好覆盖坍缩,提出her_mix方法恢复性能并大幅降低弃权偏好质量。

AI中文摘要:

事后重标记(hindsight relabeling)通过将转移的目标追溯性地替换为智能体实际达成的结果,是提高强化学习(RL)样本效率的有效工具。将其自然扩展到偏好条件多目标强化学习(MORL)时,会用智能体实际达成的偏好方向而非请求的偏好方向来重标记转移。我们表明这种扩展经常是有害的:在连续控制MO-Gymnasium套件上,跨越两种评论家骨干网络和两种偏好采样方案的四种偏好条件离策略算法中,它使36种算法-环境设置中的19种性能下降,最多达四个标准差,仅改善一种,其余不受影响。这种危害并非噪声重标记的症状;对目标去噪几乎无法恢复性能,优先采样或任何缓冲区结构选择也无法重现该危害。相反,重复重标记使评论家的覆盖范围坍缩到智能体恰好访问过的偏好空间的狭窄区域。我们将此失败模式命名为“偏好覆盖坍缩”(Preference Coverage Collapse),并用弃权偏好质量(APM)这一价值感知统计量进行量化,该统计量能追踪危害(ρ= -0.73),而纯粹的结构性覆盖计数则不能。随后我们引入her_mix,一种单参数凸组合,将达成的方向拉回请求的偏好。在每种算法和环境的一个固定值下,它使19个受损设置中的16个恢复到基线水平,保留甚至改善了重标记有帮助的一个设置,并将弃权偏好质量从69%降至6%。保护偏好单纯形上的覆盖范围,而非过滤噪声重标记,才是使事后重标记对MORL安全的关键。

英文摘要:

Hindsight relabeling which retroactively replacing a transition's goal with the outcome the agent actually achieved is an effective tool for improving sample-efficiency in Reinforcement Learning (RL). A natural extension to preference-conditioned multi-objective RL (MORL) relabels transitions with the preference direction the agent achieved rather than the one asked for. We show that this extension is frequently harmful: across four preference-conditioned off-policy algorithms spanning two critic backbones and two preference-sampling schemes on the continuous-control MO-Gymnasium suite, it degrades 19 of 36 algorithm-environment settings by as much as four standard deviations, improves only one, and leaves the rest unaffected. The harm is not a symptom of noisy relabels; denoising the target recovers almost nothing, and neither prioritized sampling nor any buffer-structural choice reproduces it. Instead, repeated relabeling collapses the critic's coverage onto whatever narrow region of the preference space the agent happened to visit. We name this failure mode \emph{Preference Coverage Collapse}, and quantify it with abandoned preference mass (APM), a value-aware statistic that tracks the harm ($ρ= -0.73$) where a purely structural coverage count does not. We then introduce \texttt{her\_mix}, a single-parameter convex combination pulling the achieved direction back towards the requested preference. At one fixed value across every algorithm and environment, it returns 16 of the 19 harmed settings to baseline, preserves and even improves the one setting in which relabeling helps, and cuts abandoned preference mass from $69\%$ to $6\%$. Protecting coverage over the preference simplex, not filtering noisy relabels, is what makes hindsight relabeling safe for MORL.

↑