arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-19 至 2025-11-19 共收录 8 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 8 篇

2511.14017 2025-11-19 cs.LG cs.AI 73%

From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs

Erum Mushtaq, Anil Ramakrishna, Satyapriya Krishna, Sattvik Sahai, Prasoon Goyal, Kai-Wei Chang, Tao Zhang, Rahul Gupta

机构 * University of Southern California(南加州大学) Amazon AGI(亚马逊人工智能实验室)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03858 2025-11-19 cs.AI cs.ET cs.MA 70%

MI9: An Integrated Runtime Governance Framework for Agentic AI

Charles L. Wang, Trisha Singhal, Ameya Kelkar, Jason Tuo

机构 * Barclays, Model Risk Management(巴克莱银行,模型风险管理部门) Columbia University(哥伦比亚大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20606 2025-11-19 cs.CL 70%

Model Editing as a Double-Edged Sword: Steering Agent Ethical Behavior Toward Beneficence or Harm

Baixiang Huang, Zhen Tan, Haoran Wang, Zijie Liu, Dawei Li, Ali Payani, Huan Liu, Tianlong Chen, Kai Shu

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

Comments AAAI 2026 Oral. 14 pages (including appendix), 11 figures. Code, data, results, and additional resources are available at: https://model-editing.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10704 2025-11-19 cs.AI 70%

The Second Law of Intelligence: Controlling Ethical Entropy in Autonomous Systems

Samih Fadli

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

Comments 12 pages, 4 figures, 1 table, includes Supplementary Materials, simulation code on GitHub (https://github.com/AerisSpace/SecondLawIntelligence )

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14606 2025-11-19 cs.CL cs.LG 62%

Bridging Human and Model Perspectives: A Comparative Analysis of Political Bias Detection in News Media Using Large Language Models

Shreya Adrita Banik, Niaz Nafi Rahman, Tahsina Moiukh, Farig Sadeque

机构 * Department of Computer Science and Engineering, BRAC University(计算机科学与工程系,BRAC大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18708 2025-11-19 cs.MA cs.AI cs.LG 62%

Skill-Aligned Fairness in Multi-Agent Learning for Collaboration in Healthcare

Promise Osaine Ekpo, Brian La, Thomas Wiener, Saesha Agarwal, Arshia Agrawal, Gonzalo Gonzalez-Pumariega, Lekan P. Molu, Angelique Taylor

机构 * Cornell Tech(康奈尔科技) Microsoft Research(微软研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14693 2025-11-19 cs.CL 57%

Talk, Snap, Complain: Validation-Aware Multimodal Expert Framework for Fine-Grained Customer Grievances

Rishu Kumar Singh, Navneet Shreya, Sarmistha Das, Apoorva Singh, Sriparna Saha

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments To be published in the Proceedings of the 40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026 Special Track on AI for Social Impact )

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13865 2025-11-19 econ.GN cs.AI q-fin.EC 57%

Randomized Controlled Trials for Conditional Access Optimization Agent

James Bono, Beibei Cheng, Joaquin Lozano

机构 * Microsoft Corporation(微软公司)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏