arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 4535 信号源:cs.CL, cs.AI, cs.LG

1. 后训练与偏好优化 4535 篇

2507.20067 2025-11-14 cs.AI cs.CL cs.LG 91%

PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training

Sarat Chandra Bobbili, Ujwal Dinesha, Dheeraj Narasimha, Srinivas Shakkottai

机构 * Texas A&M University(德克萨斯A&M大学) Inria(法国国家信息与自动化研究所)

专题命中 后训练与偏好优化 :LLM(title,abstract);post-training(title,abstract);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07163 2025-10-21 cs.CL cs.AI cs.LG 91%

Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning

Chongyu Fan, Jiancheng Liu, Licong Lin, Jinghan Jia, Ruiqi Zhang, Song Mei, Sijia Liu

机构 * Michigan State University(密歇根州立大学) University of California, Berkeley(加州大学伯克利分校) IBM Research(IBM研究院)

专题命中 后训练与偏好优化 :LLM(title,abstract);preference optimization(title,abstract);large language model(abstract);language model(abstract)

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22948 2025-09-30 cs.CL cs.AI cs.CY cs.LG 91%

SUV: Scalable Large Language Model Copyright Compliance with Regularized Selective Unlearning

Tianyang Xu, Xiaoze Liu, Feijie Wu, Xiaoqian Wang, Jing Gao

机构 * Purdue University(普渡大学)

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);preference optimization(abstract)

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06845 2025-07-08 cs.CL cs.AI cs.LG 91%

7B Fully Open Source Moxin-LLM/VLM -- From Pretraining to GRPO-based Reinforcement Learning Enhancement

Pu Zhao, Xuan Shen, Zhenglun Kong, Yixin Shen, Sung-En Chang, Arash Akbari, Timothy Rupprecht, Lei Lu, Enfu Nan, Changdi Yang, Yumei He, Weiyan Shi, Xingchen Xu, Yu Huang, Wei Jiang, Wei Wang, Yue Chen, Yong He, Yanzhi Wang

机构 * Northeastern University(东北大学) Harvard University(哈佛大学) Cornell University(康奈尔大学) Tulane University(路易斯安那州立大学) University of Washington(华盛顿大学) Futurewei Technologies(未来科技) AIBAO LLC

专题命中 后训练与偏好优化 :LLM(title,abstract);pretraining(title);large language model(abstract);language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05939 2025-04-03 cs.IR 91%

Direct Preference Optimization for LLM-Enhanced Recommendation Systems

Chao Sun, Yaobo Liang, Yaming Yang, Shilin Xu, Tianmeng Yang, Yunhai Tong

专题命中 后训练与偏好优化 :LLM(title,abstract);preference optimization(title,abstract);large language model(abstract);language model(abstract)

Comments This paper has been accepted to ICME 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04357 2025-02-10 cs.CL cs.AI cs.LG 91%

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs

Hao Sun, Yunyi Shen, Jean-Francois Ton, Mihaela van der Schaar

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);RLHF(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16950 2025-01-20 cs.CL cs.AI cs.LG 91%

Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulić, Anna Korhonen, Nigel Collier

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);RLHF(abstract)

Comments This paper has been accepted by COLM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.11385 2024-12-17 cs.CL cs.AI cs.LG 91%

Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models

Tom S. Juzek, Zina B. Ward

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);RLHF(abstract)

Comments 15 pages, 8 figures, The 31st International Conference on Computational Linguistics (COLING 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.00867 2024-11-08 cs.CR cs.AI cs.CL cs.LG 91%

Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes

Xiaomeng Hu, Pin-Yu Chen, Tsung-Yi Ho

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);RLHF(abstract)

Comments Accepted by NeurIPS 2024. Project page: https://huggingface.co/spaces/TrustSafeAI/GradientCuff-Jailbreak-Defense

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.00675 2024-10-08 cs.LG cs.AI cs.CL stat.ML 91%

Self-Play Preference Optimization for Language Model Alignment

Yue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji, Yiming Yang, Quanquan Gu

专题命中 后训练与偏好优化 :language model(title,abstract);preference optimization(title,abstract);LLM(abstract);RLHF(abstract)

Comments 27 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11827 2024-10-07 cs.CL cs.AI cs.LG 91%

WPO: Enhancing RLHF with Weighted Preference Optimization

Wenxuan Zhou, Ravi Agrawal, Shujian Zhang, Sathish Reddy Indurthi, Sanqiang Zhao, Kaiqiang Song, Silei Xu, Chenguang Zhu

专题命中 后训练与偏好优化 :RLHF(title,abstract);preference optimization(title,abstract);large language model(abstract);language model(abstract)

Comments EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.00307 2024-08-02 cs.LG cs.AI cs.CL 91%

ABC Align: Large Language Model Alignment for Safety & Accuracy

Gareth Seneque, Lap-Hang Ho, Ariel Kuperman, Nafise Erfanian Saeedi, Jeffrey Molendijk

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);post-training(abstract)

Comments 23 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.01648 2024-07-12 cs.AI cs.CL cs.LG 91%

Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation

Randall Balestriero, Romain Cosentino, Sarath Shekkizhar

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);RLHF(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02552 2024-07-04 cs.CL cs.AI cs.LG 91%

RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs

John Dang, Arash Ahmadian, Kelly Marchisio, Julia Kreutzer, Ahmet Üstün, Sara Hooker

专题命中 后训练与偏好优化 :preference optimization(title,abstract);RLHF(title);LLM(abstract);large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08114 2024-07-01 cs.LG cs.AI cs.CL 91%

Active Preference Learning for Large Language Models

William Muldrew, Peter Hayes, Mingtian Zhang, David Barber

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);RLHF(abstract);preference optimization(abstract)

Comments 13 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.15567 2024-06-25 cs.LG cs.AI cs.CL stat.ML 91%

SAIL: Self-Improving Efficient Online Alignment of Large Language Models

Mucong Ding, Souradip Chakraborty, Vibhu Agrawal, Zora Che, Alec Koppel, Mengdi Wang, Amrit Bedi, Furong Huang

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);RLHF(abstract)

Comments 24 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.08045 2024-06-04 cs.CL cs.AI cs.LG 91%

Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game

Pengyu Cheng, Yifan Yang, Jian Li, Yong Dai, Tianhao Hu, Peixin Cao, Nan Du, Xiaolong Li

专题命中 后训练与偏好优化 :LLM(title,abstract);preference optimization(title,abstract);large language model(abstract);language model(abstract)

Comments Accepted by ACL2024 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10949 2024-03-27 cs.CL cs.AI cs.LG 91%

SelfIE: Self-Interpretation of Large Language Model Embeddings

Haozhe Chen, Carl Vondrick, Chengzhi Mao

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);RLHF(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18225 2024-02-29 cs.CL cs.AI cs.LG 91%

CogBench: a large language model walks into a psychology lab

Julian Coda-Forno, Marcel Binz, Jane X. Wang, Eric Schulz

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);RLHF(abstract);prompting(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10683 2024-02-20 cs.CL cs.AI cs.LG 91%

Large Language Model Unlearning

Yuanshun Yao, Xiaojun Xu, Yang Liu

专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);RLHF(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05344 2023-10-10 cs.CL cs.AI cs.LG 91%

SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

Yi Dong, Zhilin Wang, Makesh Narsimhan Sreedhar, Xianchao Wu, Oleksii Kuchaiev

专题命中 后训练与偏好优化 :SFT(title,abstract);RLHF(title,abstract);large language model(abstract);language model(abstract)

Comments Findings of EMNLP 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16926 2026-08-19 cs.LG 新提交 91%

Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training

Data-DPO:面向大语言模型后训练中目标模型数据选择的直接偏好优化

Peng Sun, Yi Yang, Antong Zhang, Chunxiao Li, Yanbo Wang, Dianbo Liu, xin chen, Kai Yu, Lu Chen, Tianfan Fu

专题命中 后训练与偏好优化 :LLM(title);post-training(title);preference optimization(title);SFT(abstract,abstract_cn)

AI总结 针对现有数据选择方法忽略数据与目标模型能力分布兼容性的问题,提出Data-DPO方法,结合目标模型偏好、外部质量评分与边际多样性筛选数据,在Vision-Flan和LLaVA-CoT上性能优于基线及全数据训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26981 2026-07-30 cs.CL 新提交 91%

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

OptimismBench:语言模型判断中的偏差预测与对齐效应

Seonglae Cho, Adriano Koshiyama

机构 * Holistic AI(整体人工智能公司) University College London(伦敦大学学院)

专题命中 后训练与偏好优化 :LLM(summary_cn,abstract);language model(title,abstract);large language model(abstract);post-training(abstract)

AI总结 本研究推出OptimismBench基准,检测到多数LLM存在乐观方向偏差,且对齐会使模型概率产生偏差,下游流程会默认继承该偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20515 2026-07-24 cs.AI 新提交 91%

Reliability-Aware LLM Alignment from Inconsistent Human Feedback

从不一致的人类反馈中实现可靠性感知的大语言模型对齐

Jingyi Huang, Ruohan Zong, Yujun Feng, Liran Ma, Lanyu Shang, Yang Zhang

专题命中 后训练与偏好优化 :LLM(title,abstract);RLHF(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 研究如何解决人类反馈强化学习中人类注释不一致问题,提出可靠性引导的偏好优化框架RGPO,通过估计注释者可靠性、推断潜在真实标签及动态调整训练目标,有效减少训练数据不一致性和噪声,性能优于现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07856 2026-07-10 cs.AI 版本更新 91%

Dual-Difficulty Curriculum Learning for Direct Preference Optimization

用于直接偏好优化的双难度课程学习

Mengyang Li, Haozhan Geng, Zhong Zhang, Shuang Liu

机构 * Tianjin Key Laboratory of Wireless Mobile Communications and Power Transmission(天津无线移动通信与电力传输重点实验室) Tianjin Normal University(天津师范大学)

专题命中 后训练与偏好优化 :LLM(summary_cn,abstract);preference optimization(title,abstract);large language model(abstract);language model(abstract)

AI总结 研究针对大语言模型对齐中课程学习依赖一维难度视图的问题,提出将对齐难度重构为二维空间,开发DM-Curri-DPO框架,引入GSP-Curri-DPO分组自定进度学习框架,实验表明该方法提升了数据效率与鲁棒性,为LLM对齐建立新范式。

Comments We found a critical flaw in the prompt complexity metric, which affects the 2D curriculum grid construction and leads to potentially invalid comparisons. Since this undermines our main conclusions, we are withdrawing the paper and will revise the methodology before resubmission

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26997 2026-07-07 cs.DC cs.LG 新提交 91%

RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning

RolloutPipe: 分离式在线策略LLM强化学习中的流水线化Rollout与训练重叠

Rongjian Chen, Jianmin Hu, Kejiang Ye, Minxian Xu

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Southern University of Science and Technology(南方科技大学)

专题命中 后训练与偏好优化 :LLM(title,title_cn);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 针对分离式RLVR系统中rollout与训练阶段空闲问题,提出RolloutPipe框架,通过完整组流水线(CGP)和前沿组调度(FGD)实现训练与rollout重叠,保持在线策略正确性,减少训练等待时间30.7%-42.3%。

Comments 15 pages

Journal ref Proceedings of the 2026 International Conference on Cognitive Computing (ICCC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26671 2026-06-26 cs.AI 新提交 91%

NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research

NebulaExp-8B:基于全尺度消融研究的经验性后训练流程

Qiaobo Hao, Yangqian Wu, Shunyi Wang, Zhongjian Zhang, Ziqun Li, Yayin He, Muqing Li, Chen Zhong

机构 * ZTE NebulaL0 Post-Training Team(中兴通讯NebulaL0后训练团队)

专题命中 后训练与偏好优化 :SFT(summary_cn,abstract);post-training(title,abstract);large language model(abstract);language model(abstract)

AI总结 提出NebulaExp透明后训练流程,基于Qwen3-8B-base,通过数据清洗、三阶段SFT和GRPO RL优化指令与推理分支,并探索OPD/MOPD替代RL,在8B模型上实现性能提升。

Comments 29 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22305 2026-06-23 cs.CL 新提交 91%

Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning

以合适的速度学习:自适应数据调度改进大语言模型强化学习

Zicheng Xu, Ruixuan Zhang, Yu-Neng Chuang, Xiuyi Lou, Hoang Anh Duy Le, Oren Gal, Alexander S. Szalay, Zhaozhuo Xu, Guanchu Wang, Vladimir Braverman

机构 * Johns Hopkins University(约翰霍普金斯大学) Rice University(莱斯大学) University of Haifa(海法大学) Workato University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)

专题命中 后训练与偏好优化 :LLM(title,summary_cn);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 针对LLM强化学习后训练中均匀采样忽视数据语义和策略变化的问题,提出自适应数据调度框架ADS,通过语义聚类和策略边界样本选择实现双层级调度,在三个LLM和七个推理基准上平均准确率提升5.2%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16733 2026-06-16 cs.AI 新提交 91%

A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions

LLM策略优化的第一性原理推导:从期望奖励到GRPO及其结构扩展

Jianghan Shen, Siqi Luo, Yue Li, Jiyao Liu, Wanying Qu, Yi Zhang, Ziyan Huang, Tianbin Li, Ming Hu, Xiaohong Liu, Yirong Chen, Junjun He

机构 * Nanjing University(南京大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) Fudan University(复旦大学) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 后训练与偏好优化 :LLM(title,title_cn);language model(abstract);分类 cs.AI

AI总结 本文从第一性原理出发,基于轨迹概率和奖励两个轴,统一分析了从REINFORCE、PPO到GRPO及其变体的LLM策略优化方法,揭示了设计选择背后的原理和复合失败模式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14528 2026-06-15 cs.CL eess.AS 新提交 91%

BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM

BayLing-Duplex: 单一自回归LLM的原生全双工语音对话

Qingkai Fang, Shoutao Guo, Yang Feng

机构 * Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences (ICT/CAS)(中国科学院计算技术研究所智能信息处理重点实验室) Key Laboratory of AI Safety, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 后训练与偏好优化 :LLM(title,title_cn);language model(abstract);分类 cs.CL

AI总结 提出BayLing-Duplex,一种原生全双工语音语言模型,通过单个自回归LLM决定何时听、说和停止,无需外部VAD模块,仅用少量特殊标记实现,在少量微调数据上达到高交互成功率并提升响应质量。

Comments Code: https://github.com/BayLing-Models/BayLing-Duplex

详情

展开后加载摘要…

URL PDF HTML 收藏