arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-06 至 2025-11-06 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 2 篇

2511.01885 2025-11-06 cs.AI cs.LG q-bio.NC 81%

Mirror-Neuron Patterns in AI Alignment

Robyn Wyrick

机构 * Department of Computer Science University of Bath, United Kingdom(计算机科学系 英国巴斯大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 51 pages, Masters thesis. 10 tables, 7 figures, project data & code here: https://github.com/robynwyrick/mirror-neuron-frog-and-toad

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.17228 2025-11-06 cs.CY cs.AI 62%

Survey on AI Ethics: A Socio-technical Perspective

Dave Mbiazi, Meghana Bhange, Maryam Babaei, Ivaxi Sheth, Patrik Kenfack, Samira Ebrahimi Kahou

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments Updated to the peer-reviewed version accepted and published in Computational Intelligence, Volume 41, Issue 6 (Wiley, 2025)

Journal ref Computational Intelligence, Volume 41, Issue 6 (Wiley, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏