ABC-Align:具有自适应偏差控制的预测驱动对齐
ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control
浏览论文内容
中文总结 AI 辅助
ABC-Align利用丰富的伪标签信号降低方差,并通过基于人工标注子集的自适应偏差校正,在人类反馈稀缺时提升LLM对齐性能。
中文摘要 AI 辅助
语言模型的后训练常常受限于需要人工收集的偏好数据,这些数据昂贵且难以扩展。基于人工智能反馈的强化学习(RLAIF)风格的方法利用伪标签提供了丰富的替代方案,但引入了系统性偏差,降低了下游对齐质量。最近的通用半监督方法使用少量人工标注示例来纠正教师偏差,但在人工标注稀缺时尤其容易出现高方差。为此,我们提出了ABC-Align,利用丰富的伪标签信号来最小化方差,并基于人工标注子集应用轻量级、自适应的校正。校正强度在训练过程中通过相关偏差-方差量的插件估计自动调整。在人类反馈稀缺的RLHF、DPO和GRPO的LLM对齐中,我们通过一系列规模递增的实验实证表明,ABC-Align在性能上优于先前的半监督基线。我们的代码可在以下网址获取:https://this URL。
英文摘要
Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. Reinforcement learning from AI feedback (RLAIF) style approaches that leverage pseudo labels offer an abundant alternative but introduce systematic biases that degrade downstream alignment. Recent general-purpose semi-supervised methods correct for teacher bias using a small set of human-labeled examples, but suffer from high variance especially when human annotations are scarce. To this end, we propose ABC-Align, leveraging abundant pseudo label signal to minimize variance and applying a lightweight, adaptive correction grounded in the human-labeled subset. The correction strength is tuned automatically during training using plug-in estimates of the relevant bias--variance quantities. On LLM alignment with RLHF, DPO, and GRPO where human feedback is scarce, we empirically demonstrate that ABC-Align achieves superior performance over prior semi-supervised baselines in a series of experiments on an increasing scale. Our code is available at https://github.com/SewoongLab/abc-align .
发表机构
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。