arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3261 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3261 篇

2506.07520 2025-10-24 cs.SD cs.AI eess.AS 83%

LeVo: High-Quality Song Generation with Multi-Preference Alignment

Shun Lei, Yaoxun Xu, Zhiwei Lin, Huaicheng Zhang, Wei Tan, Hangting Chen, Jianwei Yu, Yixuan Zhang, Chenyu Yang, Haina Zhu, Shuai Wang, Zhiyong Wu, Dong Yu

机构 * Shenzhen International Graduate School, Tsinghua University, Shenzhen(清华大学深圳国际研究生院) Tencent AI Lab(腾讯AI实验室) Wuhan University(武汉大学) The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen)(香港中文大学(深圳)) X-LANCE Lab, Shanghai Jiao Tong University, Shanghai(上海交通大学X-LANCE实验室) School of Intelligence Science and Technology, Nanjing University, Suzhou, China(南京大学智能科学与技术学院)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01203 2025-10-21 cs.LG stat.ML 83%

KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity

Gholamali Aminian, Amir R. Asadi, Idan Shenfeld, Youssef Mroueh

机构 * The Alan Turing Institute(艾伦·图灵研究所) Statistical Laboratory(统计实验室) University of Cambridge(剑桥大学) Massachusetts Institute of Technology(麻省理工学院) IBM Research USA(IBM美国研究)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments Extra experiments are added in new version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13694 2025-10-16 cs.LG 83%

Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking

Yuchun Miao, Liang Ding, Sen Zhang, Rong Bao, Lefei Zhang, Dacheng Tao

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments 46 pages, 36 figures, submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12195 2025-10-15 cs.CL 83%

DPO-Tuned Large Language Models for Segmentation in Simultaneous Speech Translation

Zeyu Yang, Satoshi Nakamura

专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02850 2025-10-06 cs.AI 83%

Reward Model Routing in Alignment

Xinle Wu, Yao Lu

机构 * National University of Singapore(新加坡国立大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08541 2025-09-16 cs.CL 83%

CM-Align: Consistency-based Multilingual Alignment for Large Language Models

Xue Zhang, Yunlong Liang, Fandong Meng, Songming Zhang, Yufeng Chen, Jinan Xu, Jie Zhou

机构 * Key Laboratory of Big Data & Artificial Intelligence in Transportation, Beijing Jiaotong University, Ministry of Education(大数据与人工智能交通运输 key laboratory,北京交通大学,教育部) School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China(计算机科学与技术学院,北京交通大学,北京,中国) Pattern Recognition Center, WeChat AI, Tencent Inc, China(模式识别中心,微信AI,腾讯公司,中国)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20655 2025-09-12 cs.CV cs.CL 83%

Improving Alignment in LVLMs with Debiased Self-Judgment

Sihan Yang, Chenhang Cui, Zihao Zhao, Yiyang Zhou, Weilong Yan, Ying Wei, Huaxiu Yao

机构 * Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) UNC-Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02949 2025-09-09 cs.CL 83%

RADIANT: Retrieval AugmenteD entIty-context AligNmenT -- Introducing RAG-ability and Entity-Context Divergence

Vipula Rawte, Rajarshi Roy, Gurpreet Singh, Danush Khanna, Yaswanth Narsupalli, Basab Ghosh, Abhay Gupta, Argha Kamal Samanta, Aditya Shingote, Aadi Krishna Vikram, Vinija Jain, Aman Chadha, Amit Sheth, Amitava Das

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00309 2025-09-03 cs.CL 83%

Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models

Chen Zheng, Yiyuan Ma, Yuan Yang, Deyi Liu, Jing Liu, Zuquan Song, Yuxin Song, Cheng Ren, Hang Zhu, Xin Liu, Yiyuan Ma, Siyuan Qiao, Xun Zhou, Liang Xiang, Yonghui Wu

机构 * Peking University(北京大学)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16982 2025-08-26 cs.CL 83%

Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens

Ilias Chalkidis

机构 * Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments This is a working paper and will be updated with new information or corrections based on community feedback

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16455 2025-08-26 stat.ML cs.LG stat.ME 83%

On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization

Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, Weijie J. Su

机构 * University of Pennsylvania(宾夕法尼亚大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) National University of Singapore(新加坡国立大学) Peking University(北京大学) Joint corresponding authors(联合通讯作者)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments Accepted for publication in the Journal of the American Statistical Association

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08466 2025-08-13 cs.CL 83%

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints

Daren Yao, Jinsong Yuan, Ruike Chen

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18802 2025-07-28 cs.HC cs.AI 83%

DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition

Danqing Shi, Furui Cheng, Tino Weinkauf, Antti Oulasvirta, Mennatallah El-Assady

机构 * Aalto University(阿alto大学) ETH Zürich(苏黎世联邦理工学院) KTH Royal Institute of Technology(皇家理工学院)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.04427 2025-07-09 cs.CL 83%

One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversity

Sonia K. Murthy, Tomer Ullman, Jennifer Hu

机构 * School of Engineering and Applied Sciences, Harvard University(哈佛大学工程与应用科学学院) Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University(哈佛大学自然与人工智能研究学院) Department of Psychology, Harvard University(哈佛大学心理学系)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments 17 pages, 10 figures; updated with publishing information

Journal ref Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02255 2025-07-04 cs.IR cs.LG 83%

Listwise Preference Alignment Optimization for Tail Item Recommendation

Zihao Li, Chao Yang, Tong Zhang, Yakun Chen, Xianzhi Wang, Guandong Xu, Daoyi Dong

机构 * Australian Artificial Intelligence Institute (AAII) and School of Compute Sicence, Faculty of Engineering and Information Technology, University of Technology Sydney(澳大利亚人工智能研究所(AAII)和计算机科学学院,工程与信息技术学院,新南威尔士大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01368 2025-07-03 cs.CV cs.LG 83%

Activation Reward Models for Few-Shot Model Alignment

Tianning Chai, Chancharik Mitra, Brandon Huang, Gautam Rajendrakumar Gare, Zhiqiu Lin, Assaf Arbelle, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Deva Ramanan, Roei Herzig

机构 * University of California, Berkeley(加州大学伯克利分校) Carnegie Mellon University(卡内基梅隆大学) IBM Research(IBM研究院) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12999 2025-06-17 cs.CL 83%

POROver: Improving Safety and Reducing Overrefusal in Large Language Models with Overgeneration and Preference Optimization

Batuhan K. Karaman, Ishmam Zabir, Alon Benhaim, Vishrav Chaudhary, Mert R. Sabuncu, Xia Song

机构 * Cornell University(康奈尔大学) Microsoft(微软公司) Meta

专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09329 2025-06-12 cs.CL 83%

Towards Efficient and Effective Alignment of Large Language Models

Yuxin Jiang

机构 * Division of Emerging Interdisciplinary Areas(新兴跨学科领域 division)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12882 2025-06-09 cs.CR cs.CL cs.SE 83%

ProSec: Fortifying Code LLMs with Proactive Security Alignment

Xiangzhe Xu, Zian Su, Jinyao Guo, Kaiyuan Zhang, Zhenting Wang, Xiangyu Zhang

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL

Comments The first two authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01704 2025-06-06 cs.CL 83%

An Exploration of Self-Supervised Mutual Information Alignment for Multi-Task Settings

Soham V. Govande

机构 * Stanford University(斯坦福大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04063 2025-06-05 cs.HC cs.DC cs.LG 83%

Crowd-SFT: Crowdsourcing for LLM Alignment

Alex Sotiropoulos, Sulyab Thottungal Valapu, Linus Lei, Jared Coleman, Bhaskar Krishnamachari

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20417 2025-05-28 cs.AI 83%

SCAR: Shapley Credit Assignment for More Efficient RLHF

Meng Cao, Shuyuan Zhang, Xiao-Wen Chang, Doina Precup

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19428 2025-05-27 cs.CL 83%

Frictional Agent Alignment Framework: Slow Down and Don't Break Things

Abhijnan Nath, Carine Graff, Andrei Bachinin, Nikhil Krishnaswamy

机构 * Department of Computer Science, Colorado State University(计算机科学系,科罗拉多州立大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments 48 pages (main paper: 10 pages incl. Limitations and Acknowledgments; references: 6 pages; appendix: 32 pages), 9 figures, 12 tables, appearing in Proceedings of ACL 2025, Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18531 2025-05-27 cs.AI cs.CV 83%

Generative RLHF-V: Learning Principles from Multi-modal Human Preference

Jiayi Zhou, Jiaming Ji, Boyuan Chen, Jiapeng Sun, Wenqi Chen, Donghai Hong, Sirui Han, Yike Guo, Yaodong Yang

机构 * Peking University(北京大学) Hong Kong University of Science and Technology(香港科技大学) University College London(伦敦大学学院)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.AI

Comments 9 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04240 2025-05-27 cs.CL 83%

DiffPO: Diffusion-styled Preference Optimization for Efficient Inference-Time Alignment of Large Language Models

Ruizhe Chen, Wenhao Chai, Zhifei Yang, Xiaotian Zhang, Joey Tianyi Zhou, Tony Quek, Soujanya Poria, Zuozhu Liu

机构 * Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence(浙江医学影像人工智能重点实验室) Zhejiang University(浙江大学) SUTD(新加坡科技设计大学) Princeton University(普林斯顿大学) Peking University(北京大学) Nanyang Technological University(南洋理工大学) A*STAR Centre for Frontier AI Research(A*STAR前沿人工智能研究中心)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15610 2025-05-27 cs.LG 83%

On The Global Convergence Of Online RLHF With Neural Parametrization

Mudit Gaur, Amrit Singh Bedi, Raghu Pasupathy, Vaneet Aggarwal

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments The updated version of this paper is arXiv:2503.17644

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16714 2025-04-22 cs.CL 83%

Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment

Mingzhi Wang, Chengdong Ma, Qizhi Chen, Linjian Meng, Yang Han, Jiancong Xiao, Zhaowei Zhang, Jing Huo, Weijie J. Su, Yaodong Yang

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) China Telecom(中国电信) University of Pennsylvania(宾夕法尼亚大学)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11337 2025-04-16 cs.CL 83%

REWARD CONSISTENCY: Improving Multi-Objective Alignment from a Data-Centric Perspective

Zhihao Xu, Yongqi Tong, Xin Zhang, Jun Zhou, Xiting Wang

专题命中 偏好对齐 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04950 2025-04-08 cs.LG 83%

A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Wenyuan Xu, Xiaochen Zuo, Chao Xin, Yu Yue, Lin Yan, Yonghui Wu

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments 11oages,2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02708 2025-04-04 cs.CL 83%

The Hidden Space of Safety: Understanding Preference-Tuned LLMs in Multilingual context

Nikhil Verma, Manasa Bharadwaj

专题命中 偏好对齐 :safety(title,abstract);alignment(abstract);分类 cs.CL

Comments 14 pages, 11 Figures, 2 Tables, currently under review at ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏