arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-14 至 2025-11-14 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 3 篇

2511.09663 2025-11-14 cs.CY cs.AI cs.HC 81%

Alignment Debt: The Hidden Work of Making AI Usable

Cumi Oyemike, Elizabeth Akpan, Pierre Hervé-Berdys

机构 * YUX Design(YUX设计)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments 19 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10573 2025-11-14 cs.LG cs.AI cs.CL cs.HC cs.MA 80%

Towards Emotionally Intelligent and Responsible Reinforcement Learning

Garapati Keerthana, Manik Gupta

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05901 2025-11-14 cs.CL cs.AI 73%

Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations

Rui Yang, Matthew Yu Heng Wong, Huitao Li, Xin Li, Wentao Zhu, Jingchi Liao, Kunyu Yu, Jonathan Chong Kai Liew, Weihao Xuan, Yingjian Chen, Yuhe Ke, Jasmine Chiat Ling Ong, Douglas Teodoro, Chuan Hong, Daniel Shi Wei Ting, Nan Liu

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏