arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-04 至 2025-08-04 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 2 篇

2503.06072 2025-08-04 cs.CL cs.AI 62%

A Survey on Post-training of Large Language Models

Guiyao Tie, Zeli Zhao, Dingjie Song, Fuyang Wei, Rong Zhou, Yurou Dai, Wen Yin, Zhejian Yang, Jiangyue Yan, Yao Su, Zhenhan Dai, Yifeng Xie, Yihan Cao, Lichao Sun, Pan Zhou, Lifang He, Hechang Chen, Yu Zhang, Qingsong Wen, Tianming Liu, Neil Zhenqiang Gong, Jiliang Tang, Caiming Xiong, Heng Ji, Philip S. Yu, Jianfeng Gao

机构 * Huazhong University of Science and Technology(华中科技大学) Lehigh University(莱斯大学) The University of Hong Kong(香港大学) Jilin University(吉林大学) Southern University of Science and Technology(南方科技大学) Worcester Polytechnic Institute(沃思堡理工学院) LinkedIn Corporation(领英公司) Squirrel Ai Learning University of Georgia(佐治亚大学) Duke University(杜克大学) Michigan State University(密歇根州立大学) Salesforce Research(Salesforce研究) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Illinois at Chicago(伊利诺伊大学芝加哥分校) Microsoft Research(微软研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 87 pages, 21 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23454 2025-08-04 cs.HC cs.CY cs.ET cs.GR q-bio.NC 57%

Breaking the mould of Social Mixed Reality - State-of-the-Art and Glossary

Marta Bieńkiewicz, Julia Ayache, Panayiotis Charalambous, Cristina Becchio, Marco Corragio, Bertram Taetz, Francesco De Lellis, Antonio Grotta, Anna Server, Daniel Rammer, Richard Kulpa, Franck Multon, Azucena Garcia-Palacios, Jessica Sutherland, Kathleen Bryson, Stéphane Donikian, Didier Stricker, Benoît Bardy

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments pre-print

详情

展开后加载摘要…

URL PDF HTML 收藏