Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
通过分布鲁棒直接偏好优化实现鲁棒的大语言模型对齐
机构 * Texas A&M University(德克萨斯大学) ; California Institute of Technology(加州理工学院) ; Tencent AI Lab(腾讯AI实验室) ; Google DeepMind(谷歌DeepMind)
专题命中 后训练与偏好优化 :LLM(title,abstract);preference optimization(title,abstract);large language model(abstract);language model(abstract)
AI总结 本文提出Wasserstein DPO和KLDPO算法,通过分布鲁棒优化解决LLM与人类偏好间的分布偏移问题,提升对齐性能。
Comments Accepted to NeurIPS 2025