arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-12 至 2025-08-12 共收录 6 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 6 篇

2508.06849 2025-08-12 cs.CY cs.AI cs.HC 73%

Towards Experience-Centered AI: A Framework for Integrating Lived Experience in Design and Development

Sanjana Gautam, Mohit Chandra, Ankolika De, Tatiana Chakravorti, Girik Malik, Munmun De Choudhury

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07284 2025-08-12 cs.CL cs.AI cs.CY 67%

"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas

Junchen Ding, Penghao Jiang, Zihao Xu, Ziqi Ding, Yichen Zhu, Jiaojiao Jiang, Yuekang Li

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16663 2025-08-12 cs.LG cs.AI cs.CY cs.LO cs.SE 67%

Runtime Monitoring and Enforcement of Conditional Fairness in Generative AIs

Chih-Hong Cheng, Changshun Wu, Xingyu Zhao, Saddek Bensalem, Harald Ruess

机构 * Chalmers University of Technology, Sweden Carl von Ossietzky Universität Oldenburg, Germany Universit\'e Grenoble Alpes, France University of Warwick, United Kingdom CSX-AI, France SRI International, United States

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07673 2025-08-12 cs.AI cs.LG 62%

Ethics2vec: aligning automatic agents and human preferences

Gianluca Bontempi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07111 2025-08-12 cs.CL cs.AI 62%

Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution

Falaah Arif Khan, Nivedha Sivakumar, Yinong Oliver Wang, Katherine Metcalf, Cezanne Camacho, Barry-John Theobald, Luca Zappella, Nicholas Apostoloff

机构 * Apple(苹果公司)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16170 2025-08-12 cs.AI 57%

Learning How to Vote with Principles: Axiomatic Insights Into the Collective Decisions of Neural Networks

Levin Hornischer, Zoi Terzopoulou

机构 * Munich Center for Mathematical Philosophy, LMU Munich Munich Germany GATE, CNRS, Universit\'e Jean Monnet, Universit\'e Lumiere Lyon 2 Saint-Etienne France Munich Center for Mathematical Philosophy, LMU Munich GATE, CNRS, Universit\'e Jean Monnet, Universit\'e Lumiere Lyon 2

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 44 pages, 21 figures, 14 tables. Updated and published version

Journal ref Journal of Artificial Intelligence Research 83, Article 25 (August 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏