arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-08 至 2025-10-08 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 3 篇

2502.02444 2025-10-08 cs.CL cs.AI 73%

Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language Models

Haoran Ye, Tianze Zhang, Yuhang Xie, Liyuan Zhang, Yuanyi Ren, Xin Zhang, Guojie Song

机构 * State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(人工智能通用基础理论国家重点实验室,智能科学与技术学院,北京大学) Yuanpei College, Peking University(元培学院,北京大学) School of Psychological and Cognitive Sciences, Peking University(心理学与认知科学学院,北京大学) Key Laboratory of Machine Perception (Ministry of Education), Peking University(机器感知重点实验室(教育部),北京大学) PKU-Wuhan Institute for Artificial Intelligence(北京大学-武汉人工智能研究院)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06105 2025-10-08 cs.AI cs.CY cs.HC cs.LG 67%

Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences

Batu El, James Zou

机构 * Stanford University(斯坦福大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02563 2025-10-08 cs.LG cs.CL 62%

DynaGuard: A Dynamic Guardian Model With User-Defined Policies

Monte Hoover, Vatsal Baherwani, Neel Jain, Khalid Saifullah, Joseph Vincent, Chirag Jain, Melissa Kazemi Rad, C. Bayan Bruss, Ashwinee Panda, Tom Goldstein

机构 * University of Maryland(马里兰大学) Capital One

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.LG

Comments 22 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏