arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-29 至 2025-09-29 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 5 篇

2505.10426 2025-09-29 cs.CY cs.AI cs.HC math.HO 71%

Formalising Human-in-the-Loop: Computational Reductions, Failure Modes, and Legal-Moral Responsibility

Maurice Chiodo, Dennis Müller, Paul Siewert, Jean-Luc Wetherall, Zoya Yasmine, John Burden

机构 * Centre for the Study of Existential Risk(存在风险研究中心) University of Cambridge(剑桥大学) Institute of Mathematics Education(数学教育研究所) University of Cologne(科隆大学) Department of Computer Science and Technology(计算机科学与技术系) DeepFin Research(DeepFin研究) Faculty of Law(法学院) University of Oxford(牛津大学) Leverhulme Centre for the Future of Intelligence(未来智能中心)

专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.AI、cs.CY;trustworthy(comments);AI safety(comments)

Comments 31 pages. Keywords: Human-in-the-loop, Automated decision making system, Human oversight in sociotechnical systems, Oracle machine, AI safety, Trustworthy AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16355 2025-09-29 cs.LG cs.AI 62%

How Strategic Agents Respond: Comparing Analytical Models with LLM-Generated Responses in Strategic Classification

Tian Xie, Pavan Rauch, Xueru Zhang

机构 * The Ohio State University(俄亥俄州立大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Add GPT 5 experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17805 2025-09-29 cs.CY cs.AI 62%

Biospheric AI

Marcin Korecki

机构 * TU Delft(代尔夫特理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21542 2025-09-29 cs.HC cs.AI 57%

Psychological and behavioural responses in human-agent vs. human-human interactions: a systematic review and meta-analysis

Jianan Zhou, Fleur Corbett, Joori Byun, Talya Porat, Nejra van Zalk

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22424 2025-09-29 q-bio.OT 50%

Desiderata for a biomedical knowledge network: opportunities, challenges and future Directions

Chunlei Wu, Hongfang Liu, Jason Flannick, Mark A. Musen, Andrew I. Su, Lawrence Hunter, Thomas M. Powers, Cathy H. Wu

专题命中 AI治理与伦理 :trustworthy(abstract)

Comments 6 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏