arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-08 至 2025-08-08 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 2 篇

2508.05432 2025-08-08 cs.AI cs.CY 81%

Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI

Krzysztof Janowicz, Zilong Liu, Gengchen Mai, Zhangyu Wang, Ivan Majic, Alexandra Fortacz, Grant McKenzie, Song Gao

机构 * University of Vienna(维也纳大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Maine(缅因大学) McGill University(麦吉尔大学) University of Wisconsin(威斯康星大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09656 2025-08-08 cs.AI 79%

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives

Wei Zeng, Hengshu Zhu, Chuan Qin, Han Wu, Yihang Cheng, Sirui Zhang, Xiaowei Jin, Yinuo Shen, Zhenxing Wang, Feimin Zhong, Hui Xiong

机构 * Business School, Hunan University(湖南大学商学院) Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) University of Chinese Academy of Sciences(中国科学院大学) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) School of Business, Hunan University(湖南大学商学院) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能研究所) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology Hong Kong SAR(香港科技大学(香港)计算机科学与工程学院)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏