RLHF Fine-Tuning of LLMs for Alignment with Implicit User Feedback in Conversational Recommenders
专题命中 偏好对齐 :alignment(title,abstract);RLHF(title,abstract);分类 cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(title,abstract);RLHF(title,abstract);分类 cs.LG
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);RLHF(abstract);trustworthy(abstract)
Comments 47 pages, 18 figures, authors are listed in alphabetical order by their last names; v3 modifies minor issues
机构 * Microsoft(微软公司) ; University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Presented at ICML 2025
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
机构 * KAIST(韩国科学技术院) ; Seoul National University(首尔国立大学) ; Calvin University(凯尔文大学) ; NAVER AI LAB(NAVER AI实验室)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments Accepted to COLM 2025. Project Website: https://cupid.kixlab.org/
机构 * Xiaohongshu Inc.(小红书公司)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG
Comments Pre-print.Under Review
机构 * Tencent AI Lab(腾讯AI实验室)
专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.AI
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL