arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-03 至 2025-09-03 共收录 6 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 6 篇

2509.00673 2025-09-03 cs.CL cs.AI cs.IR 88%

Confident, Calibrated, or Complicit: Probing the Trade-offs between Safety Alignment and Ideological Bias in Language Models in Detecting Hate Speech

Sanjeeevan Selvaganapathy, Mehwish Nasim

机构 * School of Physics, Mathematics and Computing(物理、数学与计算学系) Network Analysis and Social Influence Modelling (NASIM) Lab(网络分析与社会影响建模实验室) The University of Western Australia(西澳大学)

专题命中 AI治理与伦理 :alignment(title,abstract);safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16534 2025-09-03 cs.CL cs.AI cs.CY 82%

Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs

Jonathan Rystrøm, Hannah Rose Kirk, Scott Hale

机构 * Oxford Internet Institute(牛津互联网研究所) University of Oxford(牛津大学)

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted at OMMM@RANLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02133 2025-09-03 cs.CL 74%

AMBEDKAR-A Multi-level Bias Elimination through a Decoding Approach with Knowledge Augmentation for Robust Constitutional Alignment of Language Models

Snehasis Mukhopadhyay, Aryan Kasat, Shivam Dubey, Rahul Karthikeyan, Dhruv Sood, Vinija Jain, Aman Chadha, Amitava Das

机构 * Indian Institute of Information Technology, Kalyani(印度信息技术学院,卡里尼) BITS Pilani Goa(比尔·斯图尔特学院,果阿) IIT Madras(马德拉斯理工学院) DTU(达丁理工大学) Artificial Intelligence Institute, University of South Carolina(南卡罗来纳大学人工智能研究所) Meta AI Amazon GenAI(亚马逊生成人工智能)

专题命中 AI治理与伦理 :alignment(title);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00085 2025-09-03 cs.CR cs.AI cs.CY 73%

Private, Verifiable, and Auditable AI Systems

Tobin South

机构 * MIT Media Lab(麻省理工学院媒体实验室)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01576 2025-09-03 cs.AI cs.CY cs.SY eess.SY 62%

Structured AI Decision-Making in Disaster Management

Julian Gerald Dcruz, Argyrios Zolotas, Niall Ross Greenwood, Miguel Arana-Catania

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 40 pages, 14 figures, 16 tables. To be published in Nature Scientific Reports

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02007 2025-09-03 cs.AI 57%

mFARM: Towards Multi-Faceted Fairness Assessment based on HARMs in Clinical Decision Support

Shreyash Adappanavar, Krithi Shailya, Gokul S Krishnan, Sriraam Natarajan, Balaraman Ravindran

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏