发表机构
University of California, Davis(加利福尼亚大学戴维斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对人类反馈可能被污染的问题,提出加权流式偏好CVaR RLHF算法,结合不确定性加权估计与乐观规划,在对抗翻转下实现分离的遗憾界并降低累积遗憾。
AI 中文摘要
基于人类反馈的强化学习(RLHF)从人类比较中学习,而这些比较可能被污染或故意操纵。本文研究了在对抗性偏好标签翻转下,具有静态条件风险价值(CVaR)的在线风险敏感RLHF。我们考虑加性线性奖励和固定参考协议,每个回合进行一次比较,在K个回合中最多有C个翻转标签。我们提出了加权流式偏好CVaR RLHF(WSP-CVaR-RLHF),该方法将不确定性加权奖励估计与乐观增强状态CVaR规划相结合。对于已知转移和归一化奖励,我们建立了遗憾界$\tilde{O}\big(\frac{d}{\kappa}\sqrt{\frac{K}{\alpha}}+\frac{dC}{\kappa\alpha}\big)$(忽略低阶项),其中$d$是奖励特征维度,$\alpha$是CVaR水平,$\kappa$刻画偏好链接。该界限将干净的统计成本与被污染反馈造成的惩罚分开。我们进一步将分析扩展到未知表格转移,其中进入CVaR目标的轨迹分布必须与奖励一起学习。我们使用矩形转移置信集、联合乐观规划和历史级CVaR模拟论证来解决由此产生的耦合不确定性。在四种对抗性攻击下的实验表明,WSP-CVaR-RLHF相对于其未加权的鲁棒对应方法,持续降低累积遗憾,同时保持置信集覆盖率。
英文摘要
Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies online risk-sensitive RLHF with static conditional value-at-risk (CVaR) under adversarial preference-label flips. We consider additive linear rewards and a fixed-reference protocol with one comparison per episode and at most $C$ flipped labels over $K$ episodes. We propose weighted streamed-preference CVaR RLHF (WSP-CVaR-RLHF), which combines uncertainty-weighted reward estimation with optimistic augmented-state CVaR planning. For known transitions and normalized rewards, we establish the regret bound $\widetilde{O}\left(\frac{d}κ\sqrt{\frac{K}α}+\frac{dC}{κα}\right)$ up to lower-order terms, where $d$ is the reward-feature dimension, $α$ is the CVaR level, and $κ$ characterizes the preference link. The bound separates the clean statistical cost from the penalty caused by corrupted feedback. We further extend the analysis to unknown tabular transitions, where the trajectory distribution entering the CVaR objective must be learned together with the reward. We address the resulting coupled uncertainty using rectangular transition confidence sets, joint optimistic planning, and a history-level CVaR simulation argument. Experiments under four adversarial attacks demonstrate that WSP-CVaR-RLHF consistently reduces cumulative regret relative to its unweighted robust counterpart while preserving confidence-set coverage.