arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-20 至 2025-08-20 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3 篇

2504.16438 2025-08-20 cs.LG cs.AI cs.CR cs.DC 62%

POPri: Private Federated Learning using Preference-Optimized Synthetic Data

Charlie Hou, Mei-Yu Wang, Yige Zhu, Daniel Lazar, Giulia Fanti

机构 * Pittsburgh Supercomputing Center, Pittsburgh, USA(匹兹堡超级计算中心) Department of ECE, Carnegie Mellon University, Pittsburgh, PA(电子工程系,卡内基梅隆大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG

Comments ICML 2025 camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13250 2025-08-20 cs.AI cs.CL cs.IR 62%

Explicit v.s. Implicit Memory: Exploring Multi-hop Complex Reasoning Over Personalized Information

Zeyu Zhang, Yang Zhang, Haoran Tan, Rui Li, Xu Chen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) National University of Singapore(新加坡国立大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

Comments 15 pages, 13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13189 2025-08-20 stat.ML cs.AI cs.LG 62%

Preference Models assume Proportional Hazards of Utilities

Chirag Nagpal

机构 * Meta Superintelligence Labs (MSL)(Meta超智能实验室)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏