Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
奉承、浮夸与雾:诊断和缓解偏好模型中的固有偏见
机构 * University of Pennsylvania(宾夕法尼亚大学) ; New York University(纽约大学)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL
AI总结 本研究通过反事实数据增强方法缓解偏好模型中的固有偏见,减少校准错误和偏斜差异,提升模型可靠性。
Comments Published at ICLR 2026; Code and data available at https://github.com/anirudhb123/preference-model-biases