arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3252 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3252 篇

2403.04283 2024-03-08 cs.CL cs.AI cs.LG 89%

Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy

Yu Zhu, Chuxiong Sun, Wenfei Yang, Wenqiang Wei, Bo Tang, Tianzhu Zhang, Zhiyu Li, Shifeng Zhang, Feiyu Xiong, Jie Hu, Mingchuan yang

专题命中 偏好对齐 :alignment(title,abstract);RLHF(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00295 2025-03-04 cs.CL cs.LG 89%

Robust Multi-Objective Preference Alignment with Online DPO

Raghav Gupta, Ryan Sullivan, Yunxuan Li, Samrat Phatale, Abhinav Rastogi

专题命中 偏好对齐 :alignment(title,abstract);DPO(title,abstract);分类 cs.CL、cs.LG

Comments AAAI 2025 - AI Alignment Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16553 2026-08-18 cs.CL 新提交 89%

STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

STAGE:面向多偏好大语言模型对齐的可控目标准入

Yongqi Tong, Zhenyu Zhang, Ruirui Wang, Kewei Fu, Shaoqing Lin, Sijie Dong, Jiang-Ming Yang, Xin Zhang, Jianshe Li

机构 * Ant International(蚂蚁国际)

专题命中 偏好对齐 :RLHF(summary_cn,abstract);alignment(title,abstract);分类 cs.CL

AI总结 该研究针对多偏好LLM对齐的目标准入时机问题,提出稳定性引导的STAGE控制器,通过活动集扩展与探测排序等方法,在多偏好对齐任务中取得优于基线的自动评估结果,为RLHF提供新控制变量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20515 2026-07-24 cs.AI 新提交 89%

Reliability-Aware LLM Alignment from Inconsistent Human Feedback

从不一致的人类反馈中实现可靠性感知的大语言模型对齐

Jingyi Huang, Ruohan Zong, Yujun Feng, Liran Ma, Lanyu Shang, Yang Zhang

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract,abstract_cn);DPO(abstract,abstract_cn);分类 cs.AI

AI总结 研究如何解决人类反馈强化学习中人类注释不一致问题,提出可靠性引导的偏好优化框架RGPO,通过估计注释者可靠性、推断潜在真实标签及动态调整训练目标,有效减少训练数据不一致性和噪声,性能优于现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28998 2026-07-07 cs.SE cs.AI 版本更新 89%

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation

从预训练或微调的大语言模型进行无奖励代码对齐:揭示代码生成中的权衡

Sanjeepan Sivapiran, Gias Uddin

机构 * York University Canada(加拿大约克大学)

专题命中 偏好对齐 :DPO(summary_cn,abstract);alignment(title,abstract);分类 cs.AI

AI总结 本研究通过实证分析,探讨了在代码生成任务中,对预训练或微调后的大语言模型进行无奖励对齐(DPO和BoNBoN)的效果,发现预训练到对齐的路径提升更大,但微调版本基线更高。

Journal ref The ACM International Conference on the Foundations of Software Engineering (FSE) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12796 2026-06-24 cs.HC cs.AI 89%

Maximizing the efficiency of human feedback in AI alignment: a comparative analysis

最大化人类反馈在AI对齐中的效率:比较分析

Andreas Chouliaras, Dimitris Chatzopoulos

机构 * University College Dublin, Ireland(都柏林大学学院)

专题命中 偏好对齐 :RLHF(summary_cn,abstract);alignment(title,abstract);分类 cs.AI

AI总结 本文探讨了RLHF中偏好推断的替代采样与评估策略,提出Swiss InfoGain方法在受限标注预算下表现更优,且更节省样本,提升了偏好学习的效率与鲁棒性。

Comments 17 pages, 6 figures, 6 algorithms. AICS2025

Journal ref 33rd International Conference on Artificial Intelligence and Cognitive Science (AICS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10217 2026-06-10 cs.LG cs.CR 新提交 89%

Alignment Defends LLMs from Property Inference Attacks

对齐防御LLM免受属性推断攻击

Pengrun Huang, Chhavi Yadav, Ruihan Wu, Kamalika Chaudhuri

机构 * University of California, San Diego(加州大学圣地亚哥分校) Carnegie Mellon University(卡内基梅隆大学) Simons Institute, UC Berkeley(伯克利大学Simons研究所)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract,abstract_cn);DPO(abstract,abstract_cn);分类 cs.LG

AI总结 提出基于对齐的防御方法,通过后训练调整模型输出分布,在不修改训练数据的情况下缓解属性推断攻击,并保持效用与机密性的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06797 2026-06-08 cs.CL 新提交 89%

Korean Culture into LLM Alignment: Toward Cultural Coherence

将韩国文化融入大语言模型对齐:迈向文化一致性

MinJae Jung, Minwoo Kim

机构 * SKT LG AI Research(LG人工智能研究) Kanana Team(Kanana团队)

专题命中 偏好对齐 :DPO(summary_cn,abstract);alignment(title,abstract);分类 cs.CL

AI总结 针对大语言模型的文化对齐,提出构建性定义而非仅抑制负面输出,设计基于提示的种子生成器扩展韩国危害分类,结合韩国法律、社会规范和解释惯例制定安全响应策略,通过DPO微调提升韩国文化安全率且不损害通用能力。

Comments Accepted to ICML 2026 Workshop on Culture X AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05370 2026-05-29 cs.CL 89%

Mining or Synthesis? Rethinking Exploration Efficiency in Iterative Alignment of Mathematical Reasoning

挖掘还是合成?重新思考数学推理迭代对齐中的探索效率

Jun Rao, Zixiong Yu, Xuebo Liu, Guhan Chen, Jing Li, Hejin Wang, Jiansheng Wei, Xiaojun Meng, Min Zhang

机构 * Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学深圳研究院) Huawei Large Model Data Technology Lab(华为大模型数据技术实验室) Huawei Multimodal Model Lab(华为多模态模型实验室) Department of Statistics and Data Science, Tsinghua University, Beijing, China(清华大学统计与数据科学系)

专题命中 偏好对齐 :DPO(summary_cn,abstract);alignment(title,abstract);分类 cs.CL

AI总结 针对数学推理任务中迭代DPO对齐时高N采样收益递减且引入噪声的问题,提出PACE框架,通过低预算探索与纠错合成偏好对,以约1/5计算量达到或超越高N基线性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07172 2026-05-11 cs.CL 89%

Topology-Enhanced Alignment for Large Language Models: Trajectory Topology Loss and Topological Preference Optimization

基于拓扑的大型语言模型对齐:轨迹拓扑损失与拓扑偏好优化

Yurui Pan, Ke Xu, Bo Peng

机构 * School of Computing and Intelligent Innovation, Fudan University(复旦大学计算与智能创新学院) School of Economics and Management, Tongji University(同济大学经济与管理学院) College of Information Technology, Shanghai Ocean University(上海海洋大学信息学院)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract,abstract_cn);DPO(abstract,abstract_cn);分类 cs.CL

AI总结 本文提出基于拓扑的对齐框架,通过轨迹拓扑损失和拓扑偏好优化提升大型语言模型的对齐性能,实验表明其在自动偏好指标和LLM判断评估中优于传统方法。

Comments Accepted to ACL 2026. 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19104 2026-04-20 cs.LG stat.ML 89%

Online Distributionally Robust LLM Alignment via Regression to Relative Reward

通过回归相对奖励实现在线分布鲁棒LLM对齐

Sharan Sahu, Martin T. Wells

机构 * Department of Statistics and Data Science(统计与数据科学系) Cornell University(康奈尔大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract,abstract_cn);DPO(abstract,abstract_cn);分类 cs.LG

AI总结 本文提出DRO-REBEL方法,通过类型-p Wasserstein、KL和χ²模糊集实现分布鲁棒优化,证明了在偏好转移下的参数误差界,并在多个基准测试中优于现有基线。

Comments 70 pages, 7 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18346 2024-06-27 cs.AI 89%

AI Alignment through Reinforcement Learning from Human Feedback? Contradictions and Limitations

Adam Dahlgren Lindström, Leila Methnani, Lea Krause, Petter Ericson, Íñigo Martínez de Rituerto de Troya, Dimitri Coelho Mollo, Roel Dobbe

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);safety(abstract);harmlessness(abstract)

Comments 12 pages, 1 table, to be submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00310 2026-08-10 cs.LG cs.AI 版本更新 88%

CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment

通过条件解码实现鲁棒的多模态安全性

Anurag Kumar, Raghuveer Peri, Jon Burnsky, Alexandru Nelus, Rohit Paturi, Srikanth Vishnubhotla, Yanjun Qi

机构 * The Ohio State University(俄亥俄州立大学) AWS(亚马逊云服务)

专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出CASA方法,通过内部表示预测安全令牌以提升多模态大语言模型的安全性,实验显示其在多种基准上显著降低攻击成功率,同时保持良性输入的实用性。

Comments 9 pages + Appendix section

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26094 2026-07-30 cs.LG cs.CL 新提交 88%

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback

用于人类反馈强化学习的元学习奖励塑形

Yunpeng Chu

专题命中 偏好对齐 :RLHF(summary_cn,abstract);DPO(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.LG

AI总结 MeRLa是元学习的任务感知奖励塑形框架,在RLHF训练前通过辅助任务元学习塑形函数,可提升LLaMA-3-8B在多基准上的对齐性能,降低训练不稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28440 2026-05-28 cs.CL cs.LG 88%

AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates

AdaDPO:具有平衡梯度更新的自适应直接偏好优化

Shaolong Chen, Madalina Ciobanu, Qingqing Mao, Ritankar Das

机构 * Incept Labs(Incept实验室)

专题命中 偏好对齐 :DPO(summary_cn,abstract);RLHF(abstract,abstract_cn);alignment(abstract);分类 cs.CL、cs.LG

AI总结 针对DPO中梯度不对称导致模型偏向避免不良回答而非生成优质回答的问题,提出AdaDPO算法,通过引入基于策略模型生成概率的自适应系数来平衡正负偏好梯度,在AlpacaEval 2上优于DPO并缓解长度偏差。

Comments 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.12339 2026-05-13 cs.LG cs.AI 88%

BSO: Safety Alignment Is Density Ratio Matching

BSO:安全对齐是密度比匹配

Tien-Phat Nguyen, Truong Nguyen, Thin Nguyen, Duy Minh Ho Nguyen, Ngoc-Thanh Dinh, Trung Le

机构 * Hanoi University of Science and Technology(河内科学技术大学) Deakin University(德金大学) Max Planck Research School for Intelligent Systems(马克斯·普朗克智能系统研究学校) VinUniversity(文大学) Monash University(墨尔本大学)

专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出BSO方法,通过密度比匹配实现安全对齐,简化了安全对齐流程,无需辅助模型,且能提升安全与有用性之间的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05113 2026-02-06 cs.AI cs.LG 88%

Democratic Preference Alignment via Sortition-Weighted RLHF

通过排序加权RLHF实现民主偏好对齐

Suvadip Sana, Jinzhou Wu, Martin T. Wells

机构 * cornell(康奈尔大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(title,abstract);分类 cs.AI、cs.LG

AI总结 DemPO通过算法排序机制优化偏好对齐,确保民主代表性,提升模型对代表性公众价值观的反映能力。

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17260 2026-01-27 cs.LG cs.AI 88%

The Viscosity of Logic: Phase Transitions and Hysteresis in DPO Alignment

逻辑的粘度:DPO对齐中的相变与滞后效应

Marco Pollanen

机构 * Department of Mathematics, Trent University, Peterborough, ON, Canada(数学系,特伦特大学,彼得伯勒,加拿大)

专题命中 偏好对齐 :alignment(title,abstract);DPO(title,abstract);分类 cs.AI、cs.LG

AI总结 研究发现DPO对齐中β参数的非单调性及滞后效应,揭示能力与边际的反相关性,推动能力解析评估方法的发展。

Comments 10 Pages, 5 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08842 2026-01-15 cs.CL cs.AI 88%

Resisting Correction: How RLHF Makes Language Models Ignore External Safety Signals in Natural Conversation

对抗修正:RLHF如何使语言模型在自然对话中忽视外部安全信号

Felipe Biava Cataneo

机构 * Felipe Biava Cataneo(独立研究者)

专题命中 偏好对齐 :RLHF(title,abstract);safety(title,abstract);分类 cs.CL、cs.AI

AI总结 RLHF使语言模型在自然对话中忽视外部安全信号,影响安全纠正的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08786 2025-12-17 cs.CL cs.AI 88%

A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs

对联邦RLHF中偏好聚合的系统评估:为LLM的多元化对齐

Mahmoud Srewa, Tianyu Zhao, Salma Elmalaki

机构 * Department of Electrical Engineering and Computer Science University of California, Irvine(电气工程与计算机科学系 加州大学伊文斯分校)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种自适应偏好聚合方法,通过动态调整权重提升联邦RLHF中LLM的公平性与对齐性能。

Comments This paper is accepted at the NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19672 2025-07-29 cs.AI cs.LG stat.ML 88%

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

Haoran Lu, Luyang Fang, Ruidong Zhang, Xinliang Li, Jiazhang Cai, Huimin Cheng, Lin Tang, Ziyu Liu, Zeliang Sun, Tao Wang, Yingchuan Zhang, Arif Hassan Zidan, Jinwen Xu, Jincheng Yu, Meizhi Yu, Hanqi Jiang, Xilin Gong, Weidi Luo, Bolun Sun, Yongkai Chen, Terry Ma, Shushan Wu, Yifan Zhou, Junhao Chen, Haotian Xiang, Jing Zhang, Afrar Jahin, Wei Ruan, Ke Deng, Yi Pan, Peilong Wang, Jiahui Li, Zhengliang Liu, Lu Zhang, Lin Zhao, Wei Liu, Dajiang Zhu, Xin Xing, Fei Dou, Wei Zhang, Chao Huang, Rongjie Liu, Mengrui Zhang, Yiwen Liu, Xiaoxiao Sun, Qin Lu, Zhen Xiang, Wenxuan Zhong, Tianming Liu, Ping Ma

机构 * Department of Statistics, University of Georgia(统计学系,佐治亚大学) School of Computing, University of Georgia(计算学院,佐治亚大学) Department of Biostatistics, Boston University(生物统计学系,波士顿大学) Department of Epidemiology & Biostatistics, University of Georgia(流行病学与生物统计学系,佐治亚大学) School of Computer and Cyber Sciences, Augusta University(计算机与网络科学学院,奥古斯塔大学) School of Electrical and Computer Engineering, University of Georgia(电气与计算机工程学院,佐治亚大学) Department of Statistics & Data Science, University of Arizona(统计学与数据科学系,亚利桑那大学) Kellogg School of Management, Northwestern University(凯洛格管理学院,西北大学) Department of Statistics, Harvard University(统计学系,哈佛大学) School of Computer Science, Carnegie Mellon University(计算机科学学院,卡内基梅隆大学)

专题命中 偏好对齐 :alignment(title,abstract);safety(title);DPO(abstract);分类 cs.AI、cs.LG

Comments 119 pages, 10 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18870 2024-12-24 cs.CL cs.AI 88%

More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness

Aaron J. Li, Satyapriya Krishna, Himabindu Lakkaraju

专题命中 偏好对齐 :alignment(title,abstract);RLHF(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14516 2024-12-20 cs.LG cs.CL 88%

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Teng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li, Vasant G Honavar

专题命中 偏好对齐 :alignment(title,abstract);DPO(title,abstract);分类 cs.CL、cs.LG

Comments Accepted by NeurIPS 2024 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10529 2024-12-17 cs.LG cs.CL 88%

Solving the Inverse Alignment Problem for Efficient RLHF

Shambhavi Krishna, Aishwarya Sahoo

专题命中 偏好对齐 :alignment(title,abstract);RLHF(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23726 2024-11-01 cs.AI cs.LG 88%

Towards Reliable Alignment: Uncertainty-aware RLHF

Debangshu Banerjee, Aditya Gopalan

专题命中 偏好对齐 :alignment(title,abstract);RLHF(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01967 2024-01-05 cs.CL cs.AI 88%

A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, Rada Mihalcea

专题命中 偏好对齐 :alignment(title,abstract);DPO(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.01581 2023-10-04 cs.LG cs.AI cs.CR 88%

On the Safety of Open-Sourced Large Language Models: Does Alignment Really Prevent Them From Being Misused?

Hangfan Zhang, Zhimeng Guo, Huaisheng Zhu, Bochuan Cao, Lu Lin, Jinyuan Jia, Jinghui Chen, Dinghao Wu

专题命中 偏好对齐 :alignment(title,abstract);safety(title);RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05040 2026-08-06 cs.CR 新提交 88%

Private Direct Preference Optimization for LLM Alignment

面向大语言模型对齐的私有直接偏好优化

Yangfan Jiang, Fei Wei, Ergute Bao, Xiaokui Xiao, Yaliang Li, Bolin Ding

专题命中 偏好对齐 :DPO(summary_cn,abstract);alignment(title,abstract)

AI总结 该研究针对DPO的隐私缺陷,提出PrivDPO方法,通过沿偏好轴添加校准随机性实现偏好隐私,在对齐基准和LLM上取得了更好的隐私-效用权衡。

Comments accepted for publication at CCS 2026, extended version

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.19320 2026-06-03 cs.CV cs.DB 88%

TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards

TextAlign: 基于层次化奖励的文本渲染偏好对齐

Mingxuan Cui, Jingpu Yang, Fengxian Ji, Qian Jiang, Zhecheng Shi, Jiaming Wang, Zirui Song, Fajri Koto, Xiuying Chen

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·穆萨大学人工智能学院) Chinese Academy of Sciences Institute of Automation(中国科学院自动化研究所) Northeastern University(东北大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 偏好对齐 :DPO(summary_cn,abstract);alignment(title,abstract)

AI总结 提出TextAlign框架,通过层次化视觉语言模型奖励将文本渲染错误分解为全局、单词和字形级别,并转化为标量偏好信号,利用GRPO或DPO进行后训练对齐,在不改变生成器架构下提升文本渲染准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17305 2026-05-20 cs.AI cs.CL cs.LG 88%

Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations

对比推理对齐:从隐藏表示中学习强化学习

Haozheng Luo, Yimin Wang, Jiahao Yu, Binghui Wang, Yan Chen

机构 * Northwestern University(西北大学) University of Michigan(密歇根大学) Illinois Institute of Technology(伊利诺伊理工学院)

专题命中 偏好对齐 :alignment(title,abstract);jailbreak(abstract,abstract_cn);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出了一种基于对比学习和强化学习的框架CRAFT,通过优化隐藏状态空间中的目标来提升对抗攻击的鲁棒性,核心贡献是通过隐藏空间的几何结构实现推理层面的安全对齐。

Comments International Conference on Machine Learning (ICML) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏