arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

共收录 4556 信号源:cs.CL, cs.AI, cs.LG

1. 后训练与偏好优化 4556 篇

2310.16271 2023-10-26 cs.CL cs.AI 86%

CycleAlign: Iterative Distillation from Black-box LLM to White-box Models for Better Human Alignment

Jixiang Hong, Quan Tu, Changyu Chen, Xing Gao, Ji Zhang, Rui Yan

专题命中 后训练与偏好优化 :LLM(title);large language model(abstract);language model(abstract);RLHF(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10202 2023-09-20 cs.CL cs.AI 86%

Stabilizing RLHF through Advantage Model and Selective Rehearsal

Baolin Peng, Linfeng Song, Ye Tian, Lifeng Jin, Haitao Mi, Dong Yu

专题命中 后训练与偏好优化 :RLHF(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

Comments 9 pages, working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.15075 2023-05-25 cs.CL cs.AI 86%

HuatuoGPT, towards Taming Language Model to Be a Doctor

Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu, Zhihong Chen, Jianquan Li, Guiming Chen, Xiangbo Wu, Zhiyi Zhang, Qingying Xiao, Xiang Wan, Benyou Wang, Haizhou Li

专题命中 后训练与偏好优化 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19504 2025-11-26 cs.LG stat.ML 86%

Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma

位置:完美AI对齐的复杂性——形式化RLHF三重困境

Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary

机构 * Berkeley AI Safety Initiative (BASIS), University of California, Berkeley(伯克利人工智能安全倡议(BASIS),加州大学伯克利分校) AWS Generative AI Innovation Center, Amazon Web Services(亚马逊网络服务生成式人工智能创新中心) Meta AI Stanford University(斯坦福大学) Northeastern University, Seattle, WA, USA(东北大学,西雅图,华盛顿州,美国)

专题命中 后训练与偏好优化 :RLHF(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG

AI总结 研究提出RLHF三重困境,指出在安全、公平和稳健之间存在根本性权衡,并通过复杂性分析证明实现全球代表性需要超多项式计算资源。

Comments Accepted at NeurIPS 2025 Workshop on Socially Responsible and Trustworthy Foundation Models (ResponsibleFM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.13948 2025-02-25 cs.CL 86%

Aligning Language Models Using Follow-up Likelihood as Reward Signal

Chen Zhang, Dading Chong, Feng Jiang, Chengguang Tang, Anningzhe Gao, Guohua Tang, Haizhou Li

专题命中 后训练与偏好优化 :language model(title,abstract);LLM(abstract,comments);preference optimization(abstract);分类 cs.CL

Comments Accepted by AAAI-2025, 16 pages, reward model, LLM Alignment, code repository at (https://github.com/e0397123/FLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26115 2026-07-30 cs.CR cs.AI cs.CL cs.LG 新提交 86%

GPT-Red: Automated Red Teaming via Self-Play at Scale

GPT-Red:基于大规模自博弈的自动化红队测试

Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cerón Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen

专题命中 后训练与偏好优化 :LLM(summary_cn,abstract);post-training(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究推出GPT-Red,一种基于大规模自博弈的自动化红队测试智能体,可发现新型提示注入攻击、攻破过往模型且泛化能力强,用于提升LLM鲁棒性并形成自我改进飞轮。

Comments 28 pages.13 main pages and 13 main figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18037 2026-07-07 cs.LG cs.AI cs.CL 版本更新 86%

Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards

梯度正则化减轻基于人类反馈和可验证奖励的强化学习中的奖励作弊

Johannes Ackermann, Michael Noukhovitch, Takashi Ishida, Masashi Sugiyama

机构 * The University of Tokyo(东京大学) RIKEN AIP(理化学研究所AIP) Mila(蒙特利尔大学Mila)

专题命中 后训练与偏好优化 :RLHF(abstract,abstract_cn);LLM(abstract);language model(abstract);post-training(abstract)

AI总结 研究基于人类反馈或可验证奖励的强化学习中奖励作弊问题,提出用梯度正则化使训练偏向奖励更准确区域,理论推导并实证验证,还改进方法,结果显示其比KL惩罚表现更好。

Comments Accepted at ICML 2026, 25 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16689 2026-06-23 eess.SP cs.NI 版本更新 86%

Against the Monolithic Wireless World Model: Why NextG Needs Composable and Agentic Intelligence

反对单一无线世界模型:为什么NextG需要可组合和代理智能

Aladin Djuhera, Farhan Ahmed, Vlad C. Andrei, Swanand Ravindra Kadhe, Alecio Binotto, Haris Gacanin, Holger Boche

专题命中 后训练与偏好优化 :LLM(summary_cn);large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 本文指出无线领域缺乏等同于LLM的数据基础,主张采用可组合和代理智能架构以实现可部署的AI原生网络。

Journal ref ICML 2026 Workshop on AI and ML for NextG

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08620 2026-06-09 cs.DB 新提交 86%

SPA: A SQL-Plan-Aware Reinforcement Learning Framework for Query Rewriting with LLMs

SPA:一种面向LLM查询重写的SQL计划感知强化学习框架

Xinyi Huang, Zhengjie Miao

专题命中 后训练与偏好优化 :LLM(title_cn,summary_cn)

AI总结 提出SPA框架,利用物理执行反馈训练LLM重写SQL查询,通过概率门控自适应奖励塑形和策略自改进,显著提升端到端运行时性能并减少有害重写。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07315 2026-06-08 cs.CL cs.AI cs.LG cs.SE 版本更新 86%

SWE-IF: Aligning Code Evaluation with Human Preference

SWE-IF: 使代码评估与人类偏好对齐

Ming Zhong, Xiang Zhou, Ting-Yun Chang, Qingze Wang, Nan Xu, Xiance Si, Dan Garrette, Shyam Upadhyay, Jeremiah Liu, Jiawei Han, Benoit Schillings, Jiao Sun

机构 * Google DeepMind(谷歌DeepMind)

专题命中 后训练与偏好优化 :LLM(summary_cn,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出SWE-IF基准,通过可验证指令分类法VeriCode评估代码指令遵循能力,发现指令遵循是区分LLM代码质量的关键,与功能正确性结合更能匹配人类偏好。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.04807 2026-06-04 cs.AI cs.CL cs.CY cs.LG 86%

BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization

BiasGRPO:通过组相对策略优化在高方差奖励景观中稳定偏差缓解

Saket Reddy, Ke Yang, ChengXiang Zhai

机构 * University of Illinois - Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 后训练与偏好优化 :RLHF(abstract,abstract_cn);large language model(abstract);language model(abstract);preference optimization(abstract)

AI总结 提出BiasGRPO框架,利用组相对策略优化(GRPO)通过归一化组内奖励来稳定大语言模型的社会偏差缓解,优于DPO和PPO。

Comments Accepted to Findings of the ACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10201 2026-05-26 cs.LG cs.AI cs.CL 86%

Future-KL Regularized GRPO: Process-Level Credit Assignment from $f$-Divergence Regularization

未来KL正则化GRPO:基于f-散度正则化的过程级信用分配

Jiarui Yao, Ruida Wang, Hao Bai, Tong Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 后训练与偏好优化 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 本文提出未来KL正则化策略优化(FRPO),通过因果未来正则化回报修正GRPO中局部KL损失缺失的梯度信号,在数学推理任务中提升pass@16并保持更高熵和更低策略漂移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21319 2026-05-19 cs.CL cs.AI cs.LG 86%

RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards

RLBFF:二进制灵活反馈用于连接人类反馈与可验证奖励

Zhilin Wang, Jiaqi Zeng, Olivier Delalleau, Ellie Evans, Daniel Egert, Hoo-Chang Shin, Felipe Soares, Yi Dong, Oleksii Kuchaiev

机构 * NVIDIA

专题命中 后训练与偏好优化 :LLM(abstract,abstract_cn);RLHF(abstract,abstract_cn);post-training(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 RLBFF结合人类偏好与规则验证,提升奖励模型对响应质量的精准捕捉,优于Bradley-Terry模型,在RM-Bench和JudgeBench上取得优异成绩,且支持用户自定义反馈原则。

Comments Published at ICLR 2026, 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18512 2026-04-21 cs.CV 86%

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models

S2H-DPO:面向视觉-语言模型的硬度感知偏好优化

Nitish Shukla, Surgan Jandial, Arun Ross

机构 * Michigan State University(密歇根州立大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 后训练与偏好优化 :language model(title,abstract);preference optimization(title)

AI总结 本文提出S2H-DPO框架,通过三级层次推理构建多图象偏好数据,提升多图象推理性能,同时保持单图象推理能力,推动视觉偏好对齐的前沿发展。

Journal ref Findings of the Association for Computational Linguistics: ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01483 2026-01-06 cs.CV 86%

Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimization

通过优势解耦偏好优化实现视觉语言模型的统一生成与自验证

Xinyu Qiu, Heng Jia, Zhengwen Zeng, Shuheng Shen, Changhua Meng, Yi Yang, Linchao Zhu

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) Venus Team, Ant Group(蚂蚁集团 Venus 团队)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);language model(title)

AI总结 ADPO通过统一生成与自验证的强化学习框架,提升视觉语言模型的验证性能和推理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05031 2025-12-09 cs.IR 86%

LSRP: A Leader-Subordinate Retrieval Framework for Privacy-Preserving Cloud-Device Collaboration

LSRP: 一种用于隐私保护云-设备协作的领导者-下属检索框架

Yingyi Zhang, Pengyue Jia, Xianneng Li, Derong Xu, Maolin Wang, Yichao Wang, Zhaocheng Du, Huifeng Guo, Yong Liu, Ruiming Tang, Xiangyu Zhao

专题命中 后训练与偏好优化 :LLM(abstract);large language model(abstract);language model(abstract);small language model(abstract)

AI总结 LSRP通过领导者-下属检索框架提升云-设备协作的隐私保护性能,增强云模型与设备模型的协同能力。

Comments Accepted at KDD'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04314 2025-03-26 cs.CV 86%

Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization

Zhanhao Liang, Yuhui Yuan, Shuyang Gu, Bohan Chen, Tiankai Hang, Mingxi Cheng, Ji Li, Liang Zheng

专题命中 后训练与偏好优化 :preference optimization(title,abstract);post-training(title)

Comments CVPR 2025. Project Page: https://rockeycoss.github.io/spo.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06769 2025-03-19 cs.SE 86%

Enhancing Trust in Language Model-Based Code Optimization through RLHF: A Research Design

Jingzhi Gong

专题命中 后训练与偏好优化 :language model(title,abstract);RLHF(title)

Comments Accepted by the Doctoral and Early Career Symposium (DECS) at ICSE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10400 2025-02-25 cs.CL cs.AI cs.LG 86%

Reinforcement Learning Enhanced LLMs: A Survey

Shuhe Wang, Shengyu Zhang, Jie Zhang, Runyi Hu, Xiaoya Li, Tianwei Zhang, Jiwei Li, Fei Wu, Guoyin Wang, Eduard Hovy

专题命中 后训练与偏好优化 :LLM(abstract);large language model(abstract);language model(abstract);RLHF(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09401 2025-02-12 cs.LG cs.AI cs.CL math.OC stat.ML 86%

Reinforcement Learning from Human Feedback with Active Queries

Kaixuan Ji, Jiafan He, Quanquan Gu

专题命中 后训练与偏好优化 :LLM(abstract);large language model(abstract);language model(abstract);RLHF(abstract)

Comments 28 pages, 1 figure, 4 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02900 2024-11-06 cs.LG cs.AI cs.CL 86%

Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Rafael Rafailov, Yaswanth Chittepu, Ryan Park, Harshit Sikchi, Joey Hejna, Bradley Knox, Chelsea Finn, Scott Niekum

专题命中 后训练与偏好优化 :LLM(abstract);large language model(abstract);language model(abstract);RLHF(abstract)

Comments 30 pages, 38th Conference on Neural Information Processing Systems (NeurIPS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.18571 2024-03-07 cs.LG cs.AI cs.CL stat.ML 86%

Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Haoxiang Wang, Yong Lin, Wei Xiong, Rui Yang, Shizhe Diao, Shuang Qiu, Han Zhao, Tong Zhang

专题命中 后训练与偏好优化 :LLM(abstract);large language model(abstract);language model(abstract);RLHF(abstract)

Comments The code and model are released at https://github.com/Haoxiang-Wang/directional-preference-alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10893 2024-02-19 cs.LG cs.AI cs.CL 86%

RLVF: Learning from Verbal Feedback without Overgeneralization

Moritz Stephan, Alexander Khazatsky, Eric Mitchell, Annie S Chen, Sheryl Hsu, Archit Sharma, Chelsea Finn

专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);RLHF(abstract);preference optimization(abstract)

Comments 9 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23311 2026-08-25 cs.CL 新提交 85%

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

超越稳定性-探索困境:面向大语言模型策略优化的环境正则化方法

Xianlei Zhou, Xiangdi Meng, Yu He, Tianyu Qi, Shuyan Guan, Xianli Zhang, Jian Zhang, Xin Li, Qika Lin, Jun Liu

机构 * Alibaba Group(阿里巴巴集团) Beijing Normal University(北京师范大学) National University of Singapore(新加坡国立大学)

专题命中 后训练与偏好优化 :LLM(title,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 该研究针对大语言模型策略优化的稳定性-探索困境,提出ERPO方法,通过查询KL散度正则化输入侧分布,接入现有策略优化流水线,在数学推理基准上实现了更强准确率与更稳定行为。

Comments Accepted to EMNLP 2026 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19893 2026-08-20 cs.CL 版本更新 85%

Future Policy Approximation for Offline Reinforcement Learning in LLM Reasoning

为离线强化学习未来政策近似改进数学推理

Minjae Oh, Yunho Choi, Dongmin Choi, Yohan Jo

机构 * Graduate School of Data Science, Seoul National University(首尔大学数据科学研究生院)

专题命中 后训练与偏好优化 :LLM(title);large language model(abstract);language model(abstract);post-training(abstract)

AI总结 本文提出未来政策近似方法,通过估计未来策略来优化离线强化学习中的梯度更新,提升长周期推理任务的稳定性与准确性。

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09595 2026-08-11 cs.AI 新提交 85%

From Sweep to Seam: Interleaved Cross-Block Post-Training Quantization

从扫描到缝合:交错跨块后训练量化

Achille Jacquemond, Yuma Ichikawa, Akira Sakai

机构 * Fujitsu Limited(富士通公司) RIKEN Center for AIP(理化学研究所先进智能项目中心) Tokai University(东海大学)

专题命中 后训练与偏好优化 :post-training(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 该研究针对大语言模型的分块后训练量化,提出交错跨块量化方法,通过重新处理块边界降低量化困惑度,提升了三比特及GPTQ等量化方案的性能。

Comments 34 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08224 2026-08-11 cs.LG 新提交 85%

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training

控制多样化强化微调:解耦RL后训练的共享控制瓶颈

Binwen Tan, Jingchao Wang, Dengzhe Hou, Lingyu Jiang, Zeyuan Wu, Yunhan Shen, Fangzhou Lin, Kazunori Yamada, Atsushi Koike

专题命中 后训练与偏好优化 :post-training(title,abstract);LLM(abstract_cn);分类 cs.LG

AI总结 本研究提出CD-RFT方法,定义后训练控制系数揭示RL后训练的共享控制瓶颈,通过正则化减少控制坍缩,在Qwen2.5-7B等模型上实现控制解耦与多任务能力提升,效果可迁移至Llama-3.2-3B。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07905 2026-08-11 cs.AI cs.RO 新提交 85%

GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning

GraphThink:用于长视距具身任务规划的图增强大语言模型思维

Chen Li, Sijie Cheng, Yuelin Zhang, Junxi Li, Maozhi Huang, Yang Liu, Wenbing Huang

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Tsinghua University(清华大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 后训练与偏好优化 :LLM(title,abstract);prompting(abstract);分类 cs.AI

AI总结 GraphThink框架结合任务图与场景图,通过上下文提示、迭代优化、GRPO奖励设计及事件驱动重规划,在ALFRED基准上实现最优性能,具强泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07019 2026-08-10 cs.AI 新提交 85%

ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization

ReQuant:用于训练后量化的固定网格离散细化

Yongge Ma, Guoan Wang, Feiyu Wang, Yaoming Li, Qian Zhang, Zihan Yan, Yinjun Han, Tong Yang

机构 * School of Computer Science, Peking University(北京大学计算机学院) School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) School of Physics, Peking University(北京大学物理学院) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Central Research Institute, ZTE Corporation(中兴公司中央研究院)

专题命中 后训练与偏好优化 :post-training(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 ReQuant是用于训练后量化(PTQ)的无反向传播固定网格细化方法,可作为即插即用后处理阶段,提升各类PTQ初始化器生成的量化模型性能,在低位宽下增益显著。

Comments 10 pages, 3 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01715 2026-08-04 cs.AI 版本更新 85%

Distributionally Robust Listwise Preference Optimization

分布鲁棒列表偏好优化

Xudong Wu, Jian Qian, Pangpang Liu, Vaneet Aggarwal, Jiayu Chen

机构 * The University of Hong Kong(香港大学) Yale University(耶鲁大学) Purdue University(普渡大学)

专题命中 后训练与偏好优化 :preference optimization(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI

AI总结 针对列表偏好排序标签不确定性,提出点态全变差鲁棒Plackett-Luce目标,将内层最大化简化为排序,在离线和在线设置中均具有理论保证,实验表明能保持干净标签性能并提升噪声鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏