Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
基于偏好向量的自适应有益有害对齐
机构 * National Taiwan University(国立台湾大学) ; Texas A&M University(德克萨斯A&M大学) ; Appier AI Research(Appier人工智能研究院) ; Graduate Institute of Communication Engineering, National Taiwan University(国立台湾大学通信工程研究所)
专题命中 偏好对齐 :alignment(title,abstract);harmlessness(title);RLHF(abstract);DPO(abstract)
AI总结 本文提出偏好向量框架,通过模块化方法实现细粒度的用户可控偏好调整,提升大语言模型的有益性和无害性平衡。
Comments Accepted at The 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026), Rabat, Morocco