arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-26 至 2025-08-26 共收录 8 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 8 篇

2508.16982 2025-08-26 cs.CL 83%

Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens

Ilias Chalkidis

机构 * Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments This is a working paper and will be updated with new information or corrections based on community feedback

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16455 2025-08-26 stat.ML cs.LG stat.ME 83%

On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization

Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, Weijie J. Su

机构 * University of Pennsylvania(宾夕法尼亚大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) National University of Singapore(新加坡国立大学) Peking University(北京大学) Joint corresponding authors(联合通讯作者)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments Accepted for publication in the Journal of the American Statistical Association

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17000 2025-08-26 cs.CL cs.LG 81%

KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF

Jason R Brown, Lennie Wells, Edward James Young, Sergio Bacallado

机构 * Computational and Biological Learning Group, Department of Engineering, University of Cambridge, Cambridge, UK(计算生物学学习组,工程系,剑桥大学,剑桥,英国) Department of Computer Science and Technology, University of Cambridge, Cambridge, UK(计算机科学与技术系,剑桥大学,剑桥,英国) Statistics Laboratory, Department of Pure Mathematics and Mathematical Statistics, University of Cambridge, UK(统计实验室,纯粹数学与数学统计系,剑桥大学,英国)

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17718 2025-08-26 cs.CV cs.AI 79%

Instant Preference Alignment for Text-to-Image Diffusion Models

Yang Li, Songlin Yang, Xiaoxuan Han, Wei Wang, Jing Dong, Yueming Lyu, Ziyu Xue

机构 * New Laboratory of Pattern Recognition, CASIA(模式识别新实验室,中国科学院自动化研究所) The Hong Kong University of Science and Technology(香港科技大学) Nanjing university(南京大学) Academy of Broadcasting Science, NRTA(广播科学研究院,国家广播电视总局)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

Comments 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16741 2025-08-26 cs.LG cs.AI 73%

WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning

Haosen Ge, Shuo Li, Lianghuan Huang

机构 * Wharton AI & Analytics Initiative(沃顿人工智能与分析倡议) University of Pennsylvania(宾夕法尼亚大学) Department of Computer and Information Science(计算机与信息科学系) Department of Physics and Astronomy(物理学与天文学系)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17637 2025-08-26 cs.CL cs.AI 62%

Weights-Rotated Preference Optimization for Large Language Models

Chenxu Yang, Ruipeng Jia, Mingyu Zheng, Naibin Gu, Zheng Lin, Siyuan Chen, Weichong Yin, Hua Wu, Weiping Wang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Baidu Inc.(百度公司)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15652 2025-08-26 cs.AI cs.IT cs.LG cs.MA math.IT 62%

Understanding Action Effects through Instrumental Empowerment in Multi-Agent Reinforcement Learning

Ardian Selmonaj, Miroslav Strupl, Oleg Szehr, Alessandro Antonucci

机构 * Istituto Dalle Molle di Studi sull’Intelligenza Artificiale (IDSIA), USI-SUPSI(日内瓦人工智能研究所(IDSIA)、USI-SUPSI)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

Comments European Conference on Artificial Intelligence (ECAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17703 2025-08-26 cs.CL 57%

EMPOWER: Evolutionary Medical Prompt Optimization With Reinforcement Learning

Yinda Chen, Yangfan He, Jing Yang, Dapeng Zhang, Zhenlong Yuan, Muhammad Attique Khan, Jamel Baili, Por Lip Yee

机构 * MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知国家重点实验室,中国科学技术大学) Department of Computer Science, University of Minnesota-Twin Cities(计算机科学系,明尼苏达大学双城分校) Center of Research for Cyber Security and Network (CSNET), Faculty of Computer Science and Information Technology, Universiti Malaya(网络安全与网络研究中心(CSNET),马来亚大学计算机科学与信息技术学院) DSLAB, School of Information Science & Engineering, Lanzhou University(信息科学与工程学院,兰州大学) Institute of Computing Technology, Chinese Academy of Sciences(计算技术研究所,中国科学院) Department of AI, Prince Mohammad bin Fahd University(人工智能系,普林姆·法赫德大学) Department of Computer Engineering, College of Computer Science, King Khalid University(计算机工程系,国王·卡利德大学)

专题命中 偏好对齐 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏