arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-16 至 2025-10-16 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 7 篇

2510.13512 2025-10-16 cs.LG cs.AI 84%

Offline and Online KL-Regularized RLHF under Differential Privacy

Yulian Wu, Rushil Thareja, Praneeth Vepakomma, Francesco Orabona

机构 * King Abdullah University of Science and Technology (KAUST)(卡布尔大学科学与技术学院) Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学) Massachusetts Institute of Technology (MIT)(麻省理工学院)

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13694 2025-10-16 cs.LG 83%

Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking

Yuchun Miao, Liang Ding, Sen Zhang, Rong Bao, Lefei Zhang, Dacheng Tao

专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.LG

Comments 46 pages, 36 figures, submitted to IEEE Transactions on Pattern Analysis and Machine Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03459 2025-10-16 cs.LG 74%

Can DPO Learn Diverse Human Values? A Theoretical Scaling Law

Shawn Im, Sharon Li

机构 * Department of Computer Sciences University of Wisconsin-Madison(计算机科学系威斯康星大学麦迪逊分校)

专题命中 偏好对齐 :DPO(title);分类 cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12041 2025-10-16 cs.CL 70%

Improving Text-to-Image Generation with Input-Side Inference-Time Scaling

Ruibo Chen, Jiacheng Pan, Heng Huang, Zhenheng Yang

机构 * TikTok University of Maryland, College Park(马里兰大学)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13501 2025-10-16 cs.AI 57%

Confidence as a Reward: Transforming LLMs into Reward Models

He Du, Bowen Li, Chengxing Xie, Chang Gao, Kai Chen, Dacheng Tao

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) Xidian University(西安电子科技大学) The Chinese University of Hong Kong(香港中文大学) Nanyang Technological University(南洋理工大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23144 2025-10-16 cs.AI cond-mat.stat-mech cs.MA nlin.AO physics.soc-ph 57%

Coordination Requires Simplification: Thermodynamic Bounds on Multi-Objective Compromise in Natural and Artificial Intelligence

Atma Anand

机构 * Department of Physics and Astronomy, University of Rochester(物理与天文学系,罗切斯特大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

Comments 15 pages, 1 figure, 9 pages supplementary material, submitted to Journal of Physics: Complexity

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10013 2025-10-16 cs.CV cs.CL 57%

Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect

Tom Kouwenhoven, Kiana Shahrasbi, Tessa Verhoef

机构 * Leiden Institute of Advanced Computer Science(莱顿先进计算机科学研究所) Leiden University(莱顿大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Presented at the Thirty-Ninth Annual Conference on Neural Information Processing Systems (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏