从不一致的人类反馈中实现可靠性感知的大语言模型对齐
Reliability-Aware LLM Alignment from Inconsistent Human Feedback
浏览论文内容
中文总结 AI 辅助
研究如何解决人类反馈强化学习中人类注释不一致问题,提出可靠性引导的偏好优化框架RGPO,通过估计注释者可靠性、推断潜在真实标签及动态调整训练目标,有效减少训练数据不一致性和噪声,性能优于现有基线。
中文摘要 AI 辅助
人类反馈强化学习(RLHF)对于使大语言模型(LLMs)与人类偏好对齐至关重要。然而,人类注释的内在不一致性和主观性常常损害其效果。现有的偏好优化框架,如直接偏好优化(DPO),通常对注释者高度分歧的模糊对与一致同意的对一视同仁,导致模型过度拟合不一致的监督信号,进而导致次优对齐。在这项工作中,我们提出了可靠性引导的偏好优化(RGPO),这是一个旨在减轻不一致人类反馈影响的稳健框架。RGPO估计注释者可靠性,并从有噪声的人类反馈中推断潜在的真实标签以识别稳健偏好。此外,我们引入了一种可靠性感知的一致性优化,根据注释的一致程度动态调整训练目标,确保模型优先处理高一致性的监督信号。在LLM对齐基准上的大量实验表明,RGPO有效地减少了训练数据中的不一致性和噪声,并且与广泛采用的RLHF基线相比实现了卓越的性能。我们的代码和配置可在这个https网址获取。
英文摘要
Reinforcement Learning from Human Feedback (RLHF) is critical for aligning Large Language Models (LLMs) with human preferences. However, its efficacy is often compromised by the inherent inconsistency and subjectivity of human annotations. Existing preference optimization frameworks, such as Direct Preference Optimization (DPO), typically treat ambiguous pairs with high annotator disagreement identically to those with unanimous consensus, forcing models to overfit to inconsistent supervision signals and leading to suboptimal alignment. In this work, we propose Reliability-Guided Preference Optimization (RGPO), a robust framework designed to mitigate the impact of inconsistent human feedback. RGPO estimates annotator reliability and infers latent ground truth labels from noisy human feedback to identify robust preferences. Furthermore, we introduce a reliability-aware consistency optimization that dynamically modulates the training objective based on the consensus level of annotations, ensuring the model prioritizes high-consensus supervision signals. Extensive experiments on LLM alignment benchmarks demonstrate that RGPO effectively reduces inconsistency and noise in training data and achieves superior performance compared to widely adopted RLHF baselines. Our code and configurations are available at https://github.com/GenieHuang/RGPO.