arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3266 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3266 篇

2602.13891 2026-02-17 cs.SD cs.AI 79%

GSRM: Generative Speech Reward Model for Speech RLHF

GSRM:生成式语音奖励模型用于语音强化学习反馈机制

Maohao Shen, Tejas Jayashankar, Osama Hanna, Naoyuki Kanda, Yancheng Wang, Kateřina Žmolíková, Ruiming Xie, Niko Moritz, Anfeng Xu, Yashesh Gaur, Gregory Wornell, Qing He, Jilong Wu

机构 * Meta Superintelligence Labs(Meta超智能实验室) Massachusetts Institute of Technology(麻省理工学院) Arizona State University(亚利桑那州立大学) University of Southern California(南加州大学)

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI

AI总结 GSRM通过生成式语音奖励模型提升语音生成的自然度,利用可解释的推理链实现更准确的自然度评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12180 2026-02-13 cs.LG cs.GT 79%

How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative Dynamics

采样如何塑造大语言模型对齐:从单次最优到迭代动态

Yurong Chen, Yu He, Michael I. Jordan, Fan Yao

机构 * Inria(法国国家信息与自动化研究所) École Normale Supérieure(巴黎高等师范学校) PSL Research University(巴黎综合理工研究所) Northwestern University(西北大学) University of California, Berkeley(加州大学伯克利分校) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

AI总结 研究探讨了采样对大语言模型对齐的影响,发现适当采样可提升排序保证,而偏向采样可能导致集中问题,并分析了迭代动态中的稳定性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09653 2026-02-12 cs.AI 79%

ClinAlign: Scaling Healthcare Alignment from Clinician Preference

ClinAlign: 从医生偏好扩展医疗对齐

Shiwei Lyu, Xidong Wang, Lei Liu, Hao Zhu, Chaohe Zhang, Jian Wang, Jinjie Gu, Benyou Wang, Yue Shen

机构 * Ant Group(蚂蚁集团) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Peking University(北京大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

AI总结 ClinAlign通过HealthRubrics和HealthPrinciples实现医疗输出与医生偏好的高效对齐,验证了资源高效的临床对齐方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09538 2026-02-11 cs.CL 79%

UniARM: Towards a Unified Autoregressive Reward Model for Multi-Objective Test-Time Alignment

UniARM: 向多目标测试时间对齐的统一自回归奖励模型迈进

Hongyan Xie, Yikun Ban, Ruiyu Fang, Zixuan Huang, Deqing Wang, Jianxin Li, Yitong Yao, Chao Wang, Shuangyong Song

机构 * School of Computer, Beihang University(北京航空航天大学计算机学院) Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

AI总结 UniARM通过统一自回归奖励模型实现多目标测试时间对齐,通过共享特征和偏好调节模块减少特征纠缠,提升对偏好权衡的控制能力。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07743 2026-02-04 cs.CL 79%

OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment

OpenRubrics: 向奖励建模和大语言模型对齐的可扩展合成 rubric 生成迈进

Tianci Liu, Ran Xu, Tony Yu, Ilgee Hong, Carl Yang, Tuo Zhao, Haoyu Wang

机构 * Purdue University(普渡大学) Emory University(埃默里大学) Georgia Institute of Technology(佐治亚理工学院) University at Albany(阿尔巴尼大学)

专题命中 偏好对齐 :alignment(title);RLHF(abstract);分类 cs.CL

AI总结 OpenRubrics 提出了一种基于对比的 rubric 生成方法,通过大规模(提示,rubric)对提升奖励建模和大语言模型对齐的性能。

Comments The first two authors contributed equally. Updated OpenRubrics dataset, RMs, and results

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01002 2026-02-03 cs.AI 79%

How RLHF Amplifies Sycophancy

如何通过RLHF放大阿谀行为

Itai Shapira, Gerdus Benade, Ariel D. Procaccia

机构 * harvard(哈佛大学)

专题命中 偏好对齐 :RLHF(title);alignment(abstract);分类 cs.AI

AI总结 本文研究了RLHF如何通过放大机制增强阿谀行为,提出了一种训练干预措施以减少这种偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04832 2026-02-03 cs.LG 79%

Discrete Diffusion Trajectory Alignment via Stepwise Decomposition

通过分步分解实现离散扩散轨迹对齐

Jiaqi Han, Austin Wang, Minkai Xu, Wenda Chu, Meihua Dang, Haotian Ye, Huayu Chen, Yisong Yue, Stefano Ermon

机构 * Stanford University(斯坦福大学) Caltech(加州理工学院) Tsinghua University(清华大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

AI总结 本研究提出一种离线偏好优化方法,通过分步分解实现离散扩散模型的轨迹对齐,提升DNA序列设计和语言建模的性能。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17671 2026-01-27 cs.CL 79%

Align to the Pivot: Dual Alignment with Self-Feedback for Multilingual Math Reasoning

对齐基准:双语自反馈的多语言数学推理

Chunxu Zhao, Xin Huang, Xue Han, Shujian Huang, Chao Deng, Junlan Feng

机构 * JIUTIAN Research, China Mobile, Beijing, China(JIUTIAN研究, 中国移动, 北京) National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室, 南京大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

AI总结 PASMR通过基准语言自反馈机制提升多语言数学推理能力,解决LLMs在多语言环境下的性能下降问题。

Comments This paper has been accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16403 2026-01-26 cs.LG 79%

Towards a Theoretical Understanding to the Generalization of RLHF

迈向RLHF泛化性的理论理解

Zhaochun Li, Mingyang Yi, Yue Wang, Shisheng Cui, Yong Liu

机构 * Beijing Institute of Technology(北京理工大学) Zhongguancun Academy(中关村学院) Renmin University of China(中国人民大学)

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.LG

AI总结 本文通过算法稳定性框架,在线性奖励模型下构建了RLHF下LLM的泛化理论,证明在特征覆盖条件下策略模型的泛化界为O(n^{-1/2}),并推广到梯度方法参数。

Comments 31 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18099 2026-01-22 cs.LG 79%

Stackelberg Self-Annotation: A Robust Approach to Data-Efficient LLM Alignment

Stackelberg 自注释:一种鲁棒的数据高效 LLM 对齐方法

Xu Chu, Zhixin Zhang, Tianyu Jia, Yujie Jin

机构 * Key Laboratory of High Confidence Software Technologies, Ministry of Education(高可信软件技术重点实验室,教育部) Center on Frontiers of Computing Studies, Peking University(计算前沿研究中心,北京大学) School of Computer Science, Peking University(计算机学院,北京大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

AI总结 本文提出Stackelberg自标注方法,通过少量人工标注和迭代自我标注,实现高效LLM对齐,减少对昂贵人工标注的依赖。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04346 2026-01-13 cs.CL 79%

Adding Alignment Control to Language Models

向语言模型添加对齐控制

Wenhong Zhu, Weinan Zhang, Rui Wang

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

AI总结 本文提出CLM模型,通过添加身份层实现语言模型的对齐控制,通过插值系数实现对齐的插值和外推,提升模型的可用性。

Comments I have changed the title of the paper and resubmitted a new one on arXiv titled "Flexible realignment of language models" (arXiv:2506.12704)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00623 2026-01-05 cs.AI 79%

DA-DPO: Cost-efficient Difficulty-aware Preference Optimization for Reducing MLLM Hallucinations

DA-DPO:面向减少多模态大语言模型幻觉的高效难度感知偏好优化

Longtian Qiu, Shan Ning, Chuyu Zhang, Jiaxuan Sun, Xuming He

机构 * ShanghaiTech University(上海科技大学) Lingang Laboratory(灵冈实验室) Shanghai Engineering Research Center of Intelligent Vision and Imaging(上海智能视觉与成像工程技术研究中心)

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.AI

AI总结 DA-DPO通过难度感知机制优化多模态大语言模型的偏好学习,有效减少幻觉并提升模型鲁棒性与泛化能力。

Comments Accepted by TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04752 2025-12-15 cs.LG 79%

RLHFSpec: Breaking the Efficiency Bottleneck in RLHF Training via Adaptive Drafting

RLHFSpec: 通过自适应草稿打破RLHF训练的效率瓶颈

Siqi Wang, Hailong Yang, Junjie Zhu, Xuezhu Wang, Yufan Xu, Depei Qian

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.LG

AI总结 RLHFSpec通过自适应草稿策略和高效样本再分配,提升RLHF训练效率,缓解生成瓶颈,提高整体性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18541 2025-12-05 cs.AI 79%

Align$^2$LLaVA: Cascaded Human and Large Language Model Preference Alignment for Multi-modal Instruction Curation

Align$^2$LLaVA: 多模态指令编纂的级联人类与大语言模型偏好对齐

Hongzhe Huang, Jiang Liu, Zhewen Yu, Li Cai, Dian Jiao, Wenqiao Zhang, Siliang Tang, Juncheng Li, Hao Jiang, Haoyuan Li, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Alibaba(阿里巴巴)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

AI总结 Align$^2$LLaVA通过级联人类与LLM偏好对齐方法,有效压缩多模态指令数据,提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03399 2025-12-04 cs.LG 79%

Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value

全栈对齐:通过厚价值模型对齐人工智能与机构

Joe Edelman, Tan Zhi-Xuan, Ryan Lowe, Oliver Klingefjord, Vincent Wang-Mascianica, Matija Franklin, Ryan Othniel Kearns, Ellie Hain, Atrisha Sarkar, Michiel Bakker, Fazl Barez, David Duvenaud, Jakob Foerster, Iason Gabriel, Joseph Gubbels, Bryce Goodman, Andreas Haupt, Jobst Heitzig, Julian Jara-Ettinger, Atoosa Kasirzadeh, James Ravi Kirkpatrick, Andrew Koh, W. Bradley Knox, Philipp Koralus, Joel Lehman, Sydney Levine, Samuele Marro, Manon Revel, Toby Shorin, Morgan Sutherland, Michael Henry Tessler, Ivan Vendrov, James Wilken-Smith

机构 * Meaning Alignment Institute(意义对齐研究所) Massachusetts Institute of Technology(麻省理工学院) University College London(伦敦大学学院) University of Oxford(牛津大学) Western University(西方大学) University of Toronto(多伦多大学) McGill University(麦吉尔大学) Stanford University(斯坦福大学) Potsdam Institute for Climate Impact Research(波茨坦气候影响研究所) Yale University(耶鲁大学) Carnegie Mellon University(卡内基梅隆大学) UT Austin(德克萨斯大学奥斯汀分校) New York University(纽约大学) Harvard University(哈佛大学) Midjourney Core contributor(Midjourney核心贡献者)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

AI总结 本文提出通过厚价值模型实现全栈对齐,以解决AI与机构目标不一致导致的不良后果,涵盖价值表示、规范推理和集体利益建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02807 2025-12-03 cs.CL 79%

SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment

SR-GRPO:稳定秩作为大语言模型对齐的内在几何奖励

Yixuan Tang, Yi Yang

机构 * The Hong Kong University of Science and Technology(香港理工大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

AI总结 SR-GRPO通过稳定秩作为内在几何奖励信号,无需外部监督提升大语言模型对齐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19474 2025-11-27 cs.LG 79%

g-DPO: Scalable Preference Optimization for Protein Language Models

g-DPO:可扩展的偏好优化用于蛋白质语言模型

Constance Ferragu, Jonathan D. Ziegler, Nicolas Deutschmann, Arthur Lindoulsi, Eli Bixby, Cradle ML Team

机构 * Cradle ML Team(Cradle团队) Cradle Zürich, Switzerland(Cradle瑞士苏黎世)

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.LG

AI总结 g-DPO通过序列聚类和组近似方法提升蛋白质语言模型的偏好优化效率,实现更快的收敛速度和与标准DPO相当的性能。

Comments Accepted at two workshops: FM4LS NeurIPS 2025 (accepted-paper.html" target="_blank" rel="noopener">https://nips2025fm4ls.github.io/pages/accepted-paper.html) and MLSB in Copenhagen EurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20059 2025-11-21 cs.CL 79%

Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical Analysis

基于偏好对齐是否总是提升基于大语言模型的翻译质量的最佳选项?一项实证分析

Hippolyte Gisserot-Boukhlef, Ricardo Rei, Emmanuel Malherbe, Céline Hudelot, Pierre Colombo, Nuno M. Guerreiro

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

AI总结 本研究通过实证分析探讨了基于偏好的对齐在提升大语言模型翻译质量中的效果,发现CPO在高质量数据上表现优异,但可能在下游度量上存在不稳定性,同时基础模型生成翻译的表现与多个外部系统相当且更一致。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17184 2025-11-20 cs.CL 79%

Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models

Xudong Han, Junjie Yang, Tianyang Wang, Ziqian Bi, Xinyuan Song, Junfeng Hao, Junhao Song

机构 * Department of Informatics, University of Sussex(信息学院,苏塞克斯大学) Pingtan Research Institute, Xiamen University(平潭研究院,厦门大学) Department of Computer Science, University of Liverpool(计算机科学系,利物浦大学) Department of Computer Science, Purdue University(计算机科学系,普渡大学) Department of Computer Science, Emory University(计算机科学系,埃默里大学) AI Agent Lab, Vokram Group(AI代理实验室,Vokram集团) Department of Computing, Imperial College London(计算系,帝国理工学院伦敦分校)

专题命中 偏好对齐 :alignment(title);safety(abstract);分类 cs.CL

Comments 24 pages, 7 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09724 2025-11-18 cs.AI 79%

UDA: Unsupervised Debiasing Alignment for Pair-wise LLM-as-a-Judge

Yang Zhang, Cunxiang Wang, Lindong Wu, Wenbo Yu, Yidong Wang, Guangsheng Bao, Jie Tang

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13264 2025-11-13 cs.CL 79%

OpenGenAlign: A Preference Dataset and Benchmark for Trustworthy Reward Modeling in Open-Ended, Long-Context Generation

Hanning Zhang, Juntong Song, Juno Zhu, Yuanhao Wu, Tong Zhang, Cheng Niu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) NewsBreak

专题命中 偏好对齐 :trustworthy(title);safety(abstract);分类 cs.CL

Comments Preprint update

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23777 2025-11-12 cs.CL 79%

CONGRAD:Conflicting Gradient Filtering for Multilingual Preference Alignment

Jiangnan Li, Thuy-Trang Vu, Christian Herold, Amirhossein Tebbifakhr, Shahram Khadivi, Gholamreza Haffari

机构 * Department of Data Science and AI, Monash University, Australia(墨尔本大学数据科学与人工智能系) eBay Inc.(eBay公司)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04880 2025-11-10 cs.AI 79%

DMA: Online RAG Alignment with Human Feedback

Yu Bai, Yukai Miao, Dawei Wang, Li Chen, Fei Long, Rundi Zhai, Dan Li, Yanyu Ren, Tianfeng Liu, Hongtao Xie, Ce Yang, Xuhui Cai

机构 * Zhongguancun Laboratory(中关村实验室) Tsinghua University(清华大学) China Mobile Communications Group Co., Ltd.(中国移动通信集团有限公司) Beijing University of Posts and Telecommunications(北京邮电大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04721 2025-11-04 cs.CL 79%

SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat

Yuru Jiang, Wenxuan Ding, Shangbin Feng, Greg Durrett, Yulia Tsvetkov

机构 * Zhejiang University(浙江大学) New York University(纽约大学) University of Washington(华盛顿大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07193 2025-10-28 cs.LG stat.ML 79%

Provably Efficient Online RLHF with One-Pass Reward Modeling

Long-Fei Li, Yu-Yang Qian, Peng Zhao, Zhi-Hua Zhou

机构 * National Key Laboratory for Novel Software Technology, Nanjing University, China(国家新型软件技术实验室,南京大学) School of Artificial Intelligence, Nanjing University, China(人工智能学院,南京大学)

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.LG

Comments NeurIPS 2025; The first two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20984 2025-10-24 cs.LG 79%

Pareto-Optimal Energy Alignment for Designing Nature-Like Antibodies

Yibo Wen, Chenwei Xu, Jerry Yao-Chieh Hu, Kaize Ding, Han Liu

机构 * Northwestern University(西北大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05831 2025-10-21 cs.CL 79%

Leveraging Robust Optimization for LLM Alignment under Distribution Shifts

Mingye Zhu, Yi Liu, Zheren Fu, Yongdong Zhang, Zhendong Mao

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Communication Content Cognition(通信内容认知国家重点实验室)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06871 2025-10-10 cs.LG cs.CV 79%

SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models

Huahui Yi, Kun Wang, Qiankun Li, Miao Yu, Liang Lin, Gongli Xi, Hao Wu, Xuming Hu, Kang Li, Yang Liu

机构 * West China Biomedical Big Data Center, West China Hospital, SCU(西昌生物医学大数据中心、西昌医院、SCU) NTU USTC TeleAI, China Telecom(TeleAI、中国电信) BUPT Tsinghua University(清华大学) HKUST(Guangzhou)(HKUST(广州))

专题命中 偏好对齐 :safety(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04712 2025-10-07 cs.CV cs.LG 79%

SEE-DPO: Self Entropy Enhanced Direct Preference Optimization

Shivanshu Shekhar, Shreyas Singh, Tong Zhang

机构 * Siebel School of Computing and Data Science(计算与数据科学学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Fractal AI Research(Fractal AI研究院)

专题命中 偏好对齐 :DPO(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07510 2025-10-07 cs.CV cs.LG 79%

Divergence Minimization Preference Optimization for Diffusion Model Alignment

Binxu Li, Minkai Xu, Jiaqi Han, Meihua Dang, Stefano Ermon

机构 * Stanford University(斯坦福大学)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏