arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 4541 信号源:cs.CL, cs.AI, cs.LG

1. 后训练与偏好优化 4541 篇

2606.28998 2026-07-07 cs.SE cs.AI 版本更新 89%

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation

从预训练或微调的大语言模型进行无奖励代码对齐:揭示代码生成中的权衡

Sanjeepan Sivapiran, Gias Uddin

机构 * York University Canada(加拿大约克大学)

专题命中 后训练与偏好优化 :LLM(title,abstract);large language model(abstract);language model(abstract);preference optimization(abstract)

AI总结 本研究通过实证分析,探讨了在代码生成任务中,对预训练或微调后的大语言模型进行无奖励对齐(DPO和BoNBoN)的效果,发现预训练到对齐的路径提升更大,但微调版本基线更高。

Journal ref The ACM International Conference on the Foundations of Software Engineering (FSE) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01357 2026-06-09 cs.LG 版本更新 89%

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning

你的自对弈算法其实是一个对抗性模仿者:通过模仿学习的视角理解LLM自对弈

Shangzhe Li, Xuchao Zhang, Chetan Bansal, Weitong Zhang

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Microsoft Research(微软研究院)

专题命中 后训练与偏好优化 :LLM(title,title_cn);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 本文通过将自对弈微调建模为模型与自身参数化的正则化隐式奖励玩家之间的极小极大博弈,统一了自对弈模仿与偏好对齐,并提出了基于χ²散度的新算法,在多种语言模型微调任务上优于现有方法。

Comments 26 pages, 6 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05881 2026-06-04 cs.CL 89%

Confidence Before Answering: A Paradigm Shift for Efficient LLM Uncertainty Estimation

回答前先置信:高效LLM不确定性估计的范式转变

Changcheng Li, Jiancan Wu, Hengheng Zhang, Zhengsu Chen, Guo An, Junxiang Qiu, Xiang Wang, Qi Tian

机构 * University of Science and Technology of China(中国科学技术大学) Huawei Inc.(华为公司)

专题命中 后训练与偏好优化 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出CoCA框架,通过GRPO强化学习联合优化置信度校准与答案准确性,实现回答前输出置信度,提升不确定性估计效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07639 2026-06-03 cs.CL 89%

Letting Tutor Personas Speak Up for LLMs: Learning Steering Vectors from Dialogue via Preference Optimization

让导师角色为LLMs发声:通过偏好优化从对话中学习引导向量

Jaewook Lee, Alexander Scarlatos, Simon Woodhead, Andrew Lan

机构 * University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校) Eedi

专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文提出使用偏好优化训练引导向量,从人类导师-学生对话中提取导师角色信息,以控制大语言模型的行为,实现多样化的教学风格。

Comments Accepted to ACL 2026 BEA Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22335 2026-05-28 cs.IR cs.AI 89%

Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation

因果直接偏好优化用于分布鲁棒的生成式推荐

Chu Zhao, Enneng Yang, Jianzhe Zhao, Guibing Guo

机构 * Northeastern University, Shenyang, China(东北大学,沈阳,中国) Shenzhen Campus of Sun Yat-sen University, China(中山大学深圳校区,中国)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 针对直接偏好优化(DPO)在生成式推荐中放大环境混杂因素导致的虚假相关性问题,提出CausalDPO,通过因果不变性学习、后门调整和软聚类环境建模来提升分布外泛化性能。

Comments 22 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21883 2026-05-27 cs.CL 89%

Token-weighted Direct Preference Optimization with Attention

基于注意力的令牌加权直接偏好优化

Chengyu Huang, Zhuohang Li, Sheng-Yen Chou, Claire Cardie

机构 * Cornell University(康奈尔大学) Vanderbilt University(范德比大学)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 提出Token-weighted DPO (TwDPO)方法,利用注意力机制估计令牌权重,在不增加额外训练成本的情况下提升大语言模型与人类偏好对齐的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15605 2026-05-26 cs.LG stat.ML 89%

Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction

自回归语言模型实际上是能量模型:对下一个词元预测的预见能力的洞察

Mathieu Blondel, Michael E. Sander, Germain Vivier-Ardisson, Tianlin Liu, Vincent Roulet

机构 * Google DeepMind(谷歌DeepMind)

专题命中 后训练与偏好优化 :language model(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);post-training(abstract)

AI总结 本文通过建立自回归模型与能量模型之间的双射,揭示了自回归模型在下一个词元预测范式下具备预见能力,并提供了理论误差界。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23803 2026-05-26 cs.CR cs.AI 89%

MultiPhishGuard: An Explainable and Adaptive Multi-Agent LLM System for Phishing Email Detection

MultiPhishGuard: 一种用于钓鱼邮件检测的可解释且自适应的多智能体大语言模型系统

Yinuo Xue, Eric Spero, Meng Wai Woo, Wei Gao, Giovanni Russello

机构 * The University of Auckland(奥克兰大学)

专题命中 后训练与偏好优化 :LLM(title,summary_cn);prompting(abstract);分类 cs.AI

AI总结 提出基于LLM的多智能体框架MultiPhishGuard,通过协调文本、URL、元数据等五个专业智能体并利用PPO动态加权,结合对抗训练提升对新型钓鱼策略的鲁棒性,在公开数据集上达到97.89%准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00224 2026-05-04 cs.AI 89%

TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization

TUR-DPO:基于拓扑和不确定性的直接偏好优化

Abdulhady Abas Abdullah, Fatemeh Daneshfar, Seyedali Mirjalili, Mourad Oussalah

机构 * Artificial Intelligence and Innovation Centre, University of Kurdistan, Erbil, Iraq(人工智能与创新中心,乌尔米耶大学,伊拉克) Department of Computer Engineering, University of Kurdistan, Iran(计算机工程系,乌尔米耶大学,伊朗) Centre for Artificial Intelligence Research and Optimisation, Torrens University Australia, Brisbane, Australia(人工智能研究与优化中心,塔伦斯大学澳大利亚,布里斯班,澳大利亚) Research and Innovation Center, Obuda University, Budapest 1034, Hungary(研究与创新中心,奥布达大学,布达佩斯1034,匈牙利) Center for Machine Vision and Signal Analysis (CMVS), University of Oulu, Finland(机器视觉与信号分析中心(CMVS),奥卢大学,芬兰)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);RLHF(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 TUR-DPO通过引入轻量级推理拓扑和结合语义忠实度、效用和拓扑质量,提升偏好对齐的稳定性与鲁棒性,同时保持训练简洁性和无需在线回滚。

Comments Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20685 2026-04-23 cs.LG 89%

MGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM Alignment

MGDA-Decoupled:基于几何的多目标优化用于基于DPO的LLM对齐

Andor Vári-Kakas, Ji Won Park, Natasa Tagasovska

机构 * Prescient Design, CS CoE, Genentech | Roche(预见设计,计算机科学学院,基因泰克 | 罗氏)

专题命中 后训练与偏好优化 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文提出MGDA-Decoupled算法,通过几何方法在DPO框架内实现更公平的多目标优化,实验显示其在UltraFeedback数据集上表现最优。

Comments Accepted to the Algorithmic Fairness Across Alignment Procedures and Agentic Systems Workshop at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19406 2026-04-22 cs.CV cs.AI 89%

HP-Edit: A Human-Preference Post-Training Framework for Image Editing

HP-Edit:一种用于图像编辑的人类偏好后训练框架

Fan Li, Chonghuinan Wang, Lina Lei, Yuping Qiu, Jiaqi Xu, Jiaxiu Jiang, Xinran Qin, Zhikai Chen, Fenglong Song, Zhixin Wang, Renjing Pei, Wangmeng Zuo

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) Harbin Institute of Technology(哈尔滨工业大学) Nankai University(南开大学)

专题命中 后训练与偏好优化 :post-training(title,abstract);RLHF(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文提出HP-Edit框架和RealPref-50K数据集,通过少量人类偏好评分数据和预训练视觉大语言模型开发自动评估器,提升图像编辑模型对人类偏好的契合度。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07449 2026-04-17 cs.IR cs.AI 89%

RLPO: Residual Listwise Preference Optimization for Long-Context Review Ranking

RLPO:长上下文评论排序的残差列表偏好优化

Hao Jiang, Zhi Yang, Annan Wang, Yichi Zhang, Weisi Lin

机构 * Nanyang Technological University(南洋理工大学) Peking University(北京大学) Independent Researcher(独立研究员)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 本文提出RLPO,通过残差列表级优化提升长上下文评论排序,解决传统方法在长上下文下的效率与准确性矛盾,实验显示其在NDCG@k上优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22718 2026-02-27 cs.AI cs.DC 89%

RLHFless: Serverless Computing for Efficient RLHF

RLHFless:用于高效RLHF的无服务器计算

Rui Wei, Hanfei Yu, Shubham Jain, Yogarajan Sivakumar, Devesh Tiwari, Jian Li, Seung-Jong Park, Hao Wang

机构 * Stevens Institute of Technology Northeastern University Stony Brook University Missouri University of Science \& Technology

专题命中 后训练与偏好优化 :RLHF(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

AI总结 RLHFless通过无服务器计算环境实现高效同步RLHF训练,预计算共享前缀并采用成本感知的演员扩展策略,提升训练速度并降低计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08136 2026-02-10 cs.CV cs.AI 89%

Robustness of Vision Language Models Against Split-Image Harmful Input Attacks

视觉语言模型对分裂图像有害输入攻击的鲁棒性

Md Rafi Ur Rashid, MD Sadik Hossain Shanto, Vishnu Asutosh Dasu, Shagufta Mehnaz

机构 * Pennsylvania State University(宾夕法尼亚州立大学) Bangladesh University of Engineering and Technology(孟加拉工程与技术大学)

专题命中 后训练与偏好优化 :language model(title,abstract);instruction tuning(abstract);pretraining(abstract);RLHF(abstract)

AI总结 本研究提出分裂图像视觉陷阱攻击(SIVA),揭示视觉语言模型在面对分裂图像攻击时的安全漏洞,并通过对抗性知识蒸馏算法提升跨模型攻击效果。

Comments 22 Pages, long conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06453 2026-02-09 cs.LG 89%

On the Plasticity and Stability for Post-Training Large Language Models

关于后训练大语言模型的可塑性与稳定性

Wenwen Qiang, Ziyin Gu, Jiahuan Zhou, Jie Hu, Jingyao Wang, Changwen Zheng, Hui Xiong

机构 * Institute of Software Chinese Academy of Sciences, Beijing, China(中国科学院软件研究所) University of the Chinese Academy of Sciences, Beijing, China(中国科学院大学) Wangxuan Institute of Computer Technology, Peking University, Beijing, China(北京大学王轩计算机技术研究所) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou), China(香港科技大学(广州)人工智能研究所) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology Hong Kong SAR, China(香港科技大学(香港特别行政区)计算机科学与工程系)

专题命中 后训练与偏好优化 :large language model(title);language model(title);post-training(title);分类 cs.LG

AI总结 本文提出PCR框架,通过概率方法解决GRPO中可塑性与稳定性之间的几何冲突,提升训练稳定性与推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00709 2025-12-02 cs.AI 89%

When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF

当人类偏好翻转时:面向RLHF的实例依赖性鲁棒损失

Yifan Xu, Xichen Ye, Yifan Chen, Qiaosheng Zhang

专题命中 后训练与偏好优化 :RLHF(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文提出了一种面向RLHF的实例依赖性鲁棒损失算法,通过建模偏好翻转机制和引入实例依赖的翻转概率,提升对齐算法的鲁棒性。

Comments Accepted by AAAI-26-AIA

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09389 2025-10-08 cs.CL 89%

Measuring LLM Novelty As The Frontier Of Original And High-Quality Output

Vishakh Padmakumar, Chen Yueh-Han, Jane Pan, Valerie Chen, He He

机构 * New York University(纽约大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 后训练与偏好优化 :LLM(title,abstract);large language model(abstract);language model(abstract);post-training(abstract)

Comments Updated results with higher coverage of open-data models and better quality judgments

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06424 2025-07-30 cs.CL cs.CY 89%

Training LLM-based Tutors to Improve Student Learning Outcomes in Dialogues

Alexander Scarlatos, Naiming Liu, Jaewook Lee, Richard Baraniuk, Andrew Lan

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Rice University(Rice大学)

专题命中 后训练与偏好优化 :LLM(title,abstract);large language model(abstract);language model(abstract);preference optimization(abstract)

Comments Published in AIED 2025: The 26th International Conference on Artificial Intelligence in Education

Journal ref In Artificial Intelligence in Education. AIED 2025. Lecture Notes in Computer Science(), vol 15877. Springer, Cham

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12854 2025-07-29 cs.CL 89%

Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

Songjun Tu, Jiahao Lin, Xiangyu Tian, Qichao Zhang, Linjing Li, Yuqian Fu, Nan Xu, Wei He, Xiangyuan Lan, Dongmei Jiang, Dongbin Zhao

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Pengcheng Laboratory(鹏城实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Wenge Technology(文生科技) Fudan University(复旦大学)

专题命中 后训练与偏好优化 :LLM(title,abstract);large language model(abstract);language model(abstract);post-training(abstract)

Comments 23pages

Journal ref COLM2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04427 2025-07-09 cs.CL 89%

One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversity

Sonia K. Murthy, Tomer Ullman, Jennifer Hu

机构 * School of Engineering and Applied Sciences, Harvard University(哈佛大学工程与应用科学学院) Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究学院) Department of Psychology, Harvard University(哈佛大学心理学系)

专题命中 后训练与偏好优化 :language model(title,abstract);LLM(abstract);large language model(abstract);post-training(abstract)

Comments 17 pages, 10 figures; updated with publishing information

Journal ref Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04686 2025-06-19 cs.AI 89%

Learning Strategic Language Agents in the Werewolf Game with Iterative Latent Space Policy Optimization

Zelai Xu, Wanjun Gu, Chao Yu, Yi Wu, Yu Wang

机构 * Tsinghua University, Beijing, China(清华大学) Beijing Zhongguancun Academy, Beijing, China(北京中关村学院) Shanghai Qi Zhi Institute, Shanghai, China(上海启智研究所)

专题命中 后训练与偏好优化 :language agent(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

Comments Published in ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07492 2025-06-10 cs.LG stat.ML 89%

Explicit Preference Optimization: No Need for an Implicit Reward Model

Xiangkun Hu, Lemin Kong, Tong He, David Wipf

机构 * The Chinese University of Hong Kong(香港中文大学) Amazon Web Services(亚马逊网络服务)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

Comments arXiv admin note: substantial text overlap with arXiv:2407.09072

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07035 2025-06-10 q-bio.BM cs.AI 89%

AnnoDPO: Protein Functional Annotation Learning with Direct Preference Optimization

Zixuan Jiang, Renjing Xu

专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00888 2025-04-22 cs.CL cs.HC 89%

Aligning Language Models with Demonstrated Feedback

Omar Shaikh, Michelle S. Lam, Joey Hejna, Yijia Shao, Hyundong Cho, Michael S. Bernstein, Diyi Yang

机构 * Stanford University(斯坦福大学) USC(美国南加州大学)

专题命中 后训练与偏好优化 :language model(title,abstract);LLM(abstract);RLHF(abstract);preference optimization(abstract)

Comments ICLR 2025; 28 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04950 2025-04-08 cs.LG 89%

A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Wenyuan Xu, Xiaochen Zuo, Chao Xin, Yu Yue, Lin Yan, Yonghui Wu

专题命中 后训练与偏好优化 :RLHF(title,abstract);large language model(abstract);language model(abstract);foundation model(abstract)

Comments 11oages,2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01233 2025-03-04 cs.CL 89%

PEO: Improving Bi-Factorial Preference Alignment with Post-Training Policy Extrapolation

Yuxuan Liu

专题命中 后训练与偏好优化 :post-training(title,abstract);large language model(abstract);language model(abstract);RLHF(abstract)

Comments Technical report, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10914 2025-02-21 cs.CL 89%

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment

Sizhe Wang, Yongqi Tong, Hengyuan Zhang, Dawei Li, Xin Zhang, Tianlong Chen

专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

Comments The 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL 2025)- Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12735 2025-02-10 cs.LG 89%

Online Preference Alignment for Language Models via Count-based Exploration

Chenjia Bai, Yang Zhang, Shuang Qiu, Qiaosheng Zhang, Kang Xu, Xuelong Li

专题命中 后训练与偏好优化 :language model(title,abstract);LLM(abstract);large language model(abstract);RLHF(abstract)

Comments Accepted by ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12895 2025-01-23 cs.CL 89%

Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback

Yafu Li, Xuyang Hu, Xiaoye Qu, Linjie Li, Yu Cheng

专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

Comments 43 pages; work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.19443 2025-01-23 cs.CL 89%

Mixed Preference Optimization: Reinforcement Learning with Data Selection and Better Reference Model

Qi Gou, Cam-Tu Nguyen

专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏