Scalable Valuation of Human Feedback through Provably Robust Model Alignment
机构 * The University of Osaka(大阪大学) ; Lattice Lab, Toyota Motor Corporation(丰田公司Lattice实验室) ; Machine Learning Research Group, University of Oxford(牛津大学机器学习研究组) ; RIKEN AIP(理化学研究所AIP)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG
Comments Accepted by the 39th Conference on Neural Information Processing Systems (NeurIPS2025), 49 pages, 7 figures