arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-04 至 2025-11-04 共收录 7 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 7 篇

2511.00379 2025-11-04 cs.AI cs.CL 81%

Diverse Human Value Alignment for Large Language Models via Ethical Reasoning

Jiahao Wang, Songkai Xue, Jinghui Li, Xiaozhen Wang

专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by AIES 2025, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09606 2025-11-04 cs.CY 70%

Local US officials' views on the impacts and governance of AI: Evidence from 2022 and 2023 survey waves

Sophia Hatz, Noemi Dreksler, Kevin Wei, Baobao Zhang

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Journal ref PLoS One 20(10): e0332919, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00024 2025-11-04 cs.CY cs.AI cs.CL cs.LG stat.AP 70%

Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model

Haotian Hang, Yueyang Shen, Vicky Zhu, Jose Cruz, Michelle Li

机构 * University of Southern California(南加州大学) University of Michigan(密歇根大学) Babson College(巴布森学院) University of Connecticut(康涅狄格大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20432 2025-11-04 cs.AI cs.CY cs.GT cs.LG 67%

LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory

Jingru Jia, Zehua Yuan, Junhao Pan, Paul E. McNamara, Deming Chen

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01550 2025-11-04 cs.AI 57%

Analyzing Sustainability Messaging in Large-Scale Corporate Social Media

Ujjwal Sharma, Stevan Rudinac, Ana Mićković, Willemijn van Dolen, Marcel Worring

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14886 2025-11-04 cs.CV 50%

Surgical Scene Understanding in the Era of Foundation AI Models: A Comprehensive Review

Ufaq Khan, Umair Nawaz, Adnan Qayyum, Shazad Ashraf, Yutong Xie, Muhammad Haris Khan, Muhammad Bilal, Junaid Qadir

机构 * MBZ University of AI(MBZ人工智能大学) Hamad Bin Khalifa University(哈马德·本·卡西姆大学) Birmingham City University(伯明翰城市大学) Qatar University(卡塔尔大学) University Hospitals Birmingham and University of Birmingham(伯明翰大学医院和伯明翰大学)

专题命中 AI治理与伦理 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00026 2025-11-04 cs.RO 50%

Gen AI in Automotive: Applications, Challenges, and Opportunities with a Case study on In-Vehicle Experience

Chaitanya Shinde, Divya Garikapati

专题命中 AI治理与伦理 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏