arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BA-DPO:用于语言模型对齐的偏差调整直接偏好优化

BA-DPO: Bias-Adjusted Direct Preference Optimization for Language Model Alignment

Antonio Ferrara, Alberto Rumi, Francesco Bonchi

arXiv 2609.35044首次发表:更新:

发表机构

Intesa Sanpaolo AI Research(意大利联合圣保罗银行人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出BA-DPO,通过为每位标注者添加偏差参数来消除任意属性偏差,在植入偏差语料上消除DPO偏移的81-95%,在真实标注上消除约一半长度增加,且不增加KL或损失质量。

AI 中文摘要

基于偏好的对齐方法(如直接偏好优化(DPO))使用由人类标注者标记的成对偏好来微调语言模型。然而,标注者会对某些属性带有系统性偏差:如暗示性别或种族的名字、人物角色、语言变体、格式惯例或长度。如果处理不当,这些系统性偏差会在对齐过程中被吸收并放大。现有方法解决了长度偏差或标注者分歧问题,但未能消除对任意属性的偏差。为解决这一局限,我们提出了偏差调整DPO(BA-DPO),这是DPO的一种推广,为每位标注者针对带有声明属性的响应增加一个偏差参数。我们证明了该目标在偏差参数上是凸的,并且投票可以识别每位标注者的偏差,直至一个共享常数。剩余的常数决定了对齐模型的属性比率:默认情况下为参考模型的比率,或目标比率,我们用它来使有偏差的策略达到统计均等。在带有植入偏差的语料库上,DPO将属性从平衡起点驱动到概率0.96,而BA-DPO消除了该偏移的81%至95%;在具有真实标注者的MultiPref上,它消除了DPO长度增加的大约一半。两者在0.5B规模下使用全微调,在8B规模下使用LoRA,均保持与DPO相同的KL且不损失评判质量。

英文摘要

Preference-based alignment methods such as Direct Preference Optimization (DPO) use pairwise preferences labeled by human annotators to fine-tune language models. However, annotators carry systematic biases toward some attributes: a name that signals a gender or an ethnicity, a persona, a language variety, a formatting convention, or length. If not properly addressed, these systematic biases can be absorbed and amplified during alignment. Existing methods address length bias or annotator disagreement, but fail to eliminate biases toward arbitrary attributes. To address this limitation, we propose Bias-Adjusted DPO (BA-DPO), a generalization of DPO that adds one bias parameter per annotator toward responses carrying a declared attribute. We prove that the objective is convex in the bias parameters and that the votes identify each annotator's bias up to a shared constant. The remaining constant is what fixes the aligned model's attribute rate: by default the rate of the reference model, or a target rate, which we use to bring a biased policy to statistical parity. On a corpus with planted biases, DPO drives the attribute from a balanced start to probability 0.96 and BA-DPO removes 81 to 95\% of that shift; on MultiPref with real annotators it removes about half of DPO's lengthening. Both hold at 0.5B with full fine-tuning and at 8B with LoRA, at no higher KL than DPO and no loss in judged quality.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑