Supervised Fine-Tuning as Inverse Reinforcement Learning
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 8 pages
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments ICLR 2024. Code is available on our project website: https://xingyaoww.github.io/mint-bench
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments First two authors contributed equally; Project website: https://selma-t2i.github.io/
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 9 pages, 9 figures
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Updates for camera-ready submission
Journal ref NeurIPS Workshop on Generative AI for Education (GAIED), 2023
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 25 pages, 6 figures
专题命中 偏好对齐 :alignment(abstract);RLHF(abstract)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Accepted to EMNLP 2023 main conference
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments technical report. arXiv admin note: text overlap with arXiv:2306.16636 by other authors
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Proceedings of the 40th International Conference on Machine Learning (ICML), 2023
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Preprint. Code at https://github.com/FranxYao/chain-of-thought-hub
专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.AI、cs.LG
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG
Comments for associated data visualizations, see https://www.evals.anthropic.com/model-written/ for full datasets, see https://github.com/anthropics/evals
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments AIES 2020
可引导的文化偏好优化奖励模型
机构 * Stanford University(斯坦福大学) ; University of Amsterdam(阿姆斯特丹大学)
专题命中 偏好对齐 :alignment(abstract,comments);分类 cs.CL、cs.AI
AI总结 提出SCPO算法,通过平衡多种文化偏好训练奖励模型,在PRISM和GlobalOpinionQA数据集上提升少数群体偏好预测准确率最多7点,训练效率提高280%。
Comments Accepted to Pluralistic Alignment @ ICML 2026
机构 * Duke University(杜克大学) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 偏好对齐 :alignment(abstract,comments);分类 cs.AI、cs.CY
Comments To appear in the AAAI 2026 Alignment Track
机构 * KAIST(韩国科学技术院) ; Columbia University(哥伦比亚大学)
专题命中 偏好对齐 :RLHF(abstract,comments);分类 cs.AI、cs.LG
Comments Accepted at ACL 2025, Source code: https://github.com/mintaywon/IF_RLHF
Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 63 (2025) 27471-27500
机构 * Department of Electrical and Computer Engineering, University of Texas at San Antonio, Texas, USA(电子与计算机工程系,德克萨斯州立大学圣安东尼奥分校) ; DEVCOM Army Research Lab, USA(陆军研究实验室)
专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG;alignment(comments)
Comments Accepted to the workshop on Models of Human Feedback for AI Alignment at the 42nd International Conference on Machine Learning
专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.LG;trustworthy(comments)
Comments Accepted in TrustNLP: Third Workshop on Trustworthy Natural Language Processing, co-located with ACL 2023
EstLLM:通过持续预训练和后训练增强多语言大语言模型中的爱沙尼亚能力
机构 * Institute of Computer Science, University of Tartu(塔尔图大学计算机科学研究院) ; Department of Software Science, Tallinn University of Technology(塔林技术大学软件科学系) ; Institute of the Estonian Language, Tallinn, Estonia(爱沙尼亚语言研究院) ; School of Humanities, Tallinn University(塔林大学人文学院)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI
AI总结 EstLLM通过持续预训练和后训练提升多语言大语言模型中爱沙尼亚的能力,增强语言和推理表现。
大语言模型在不确定性下的偏好推理
机构 * Penn State University(宾夕法尼亚州立大学)
专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG
AI总结 本文聚焦大语言模型决策智能体的偏好推理问题,将不确定性挑战形式化为认知与结构两类,发现当前模型无法区分确定与不确定实例,推理校准不当。
Comments 55 pages, 14 figures
谁的“金标准”?标注者群体在条目层面分歧巨大,却被小型排行榜掩盖
专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.LG
AI总结 该研究发现标注者群体在条目层面分歧巨大,小型模型排行榜的一致性是假象,量化了其脆弱性,证明某数据集的标注者无差异假设错误,LLM评判器更贴合大众标注者群体。
Comments Submitted to the HAIC workshop at NeurIPS 2026
PEER:统一的过程-结果强化学习用于结构化共情推理
机构 * Shandong University(山东大学) ; Shandong Jianzhu University(山东建筑大学) ; Singapore Management University(新加坡管理大学)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI
AI总结 本文提出PEER,通过结构化共情推理框架提升情感支持对话的 empathy、策略一致性及人性表现,采用GRPO与UnifiReward模型,结合数据增强减少重复。
VDC-Agent:当视频详细描述器通过代理自我反思而自我进化
机构 * Xi’an Jiaotong University(西安交通大学) ; Kuaishou Technology(快手科技) ; Shenzhen University of Advanced Technology(深圳先进技术大学)
专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG
AI总结 VDC-Agent通过自我反思机制实现视频详细描述的自我进化,利用自动生成的(描述,评分)对提升描述准确性与评分表现。
Comments Accepted to ECCV 2026. Project Page: https://vdcagent.github.io