arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-11 至 2025-08-11 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 4 篇

2508.05846 2025-08-11 cs.CY cs.AI cs.HC cs.LG cs.RO 82%

Towards Transparent Ethical AI: A Roadmap for Trustworthy Robotic Systems

Ahmad Farooq, Kamran Iqbal

机构 * University of Arkansas at Little Rock(阿拉巴马州立大学)

专题命中 AI治理与伦理 :trustworthy(title,abstract);分类 cs.AI、cs.CY、cs.LG

Comments Published in the Proceedings of the 2025 3rd International Conference on Robotics, Control and Vision Engineering (RCVE'25). 6 pages, 3 tables

Journal ref RCVE'25: Proceedings of the 2025 3rd International Conference on Robotics, Control and Vision Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05938 2025-08-11 cs.CL cs.AI cs.CY 75%

Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale

Rafal Kocielnik, Min Kim, Penphob, Boonyarungsrit, Fereshteh Soltani, Deshawn Sambrano, Animashree Anandkumar, R. Michael Alvarez

机构 * California Institute of Technology(加利福尼亚理工学院) Activision Publishing, Inc.(暴雪娱乐公司)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 9 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05913 2025-08-11 cs.HC cs.AI cs.CL 62%

Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction

Stefan Pasch, Min Chul Cha

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06479 2025-08-11 cs.CY 57%

The Problem of Atypicality in LLM-Powered Psychiatry

Bosco Garcia, Eugene Y. S. Chua, Harman Singh Brah

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Preprint of 8/8/2025 -- please cite published version. This article has been published in the Journal of Medical Ethics (2025) following peer review and can also be viewed on the journal's website at 10.1136/jme-2025-110972

详情

展开后加载摘要…

URL PDF HTML 收藏