arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1852 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1852 篇

2506.12245 2025-06-17 cs.AI 57%

Reversing the Paradigm: Building AI-First Systems with Human Guidance

Cosimo Spera, Garima Agrawal

机构 * Minerva CQ

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11068 2025-06-16 cs.CL 57%

Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language Models

Bumjin Park, Jinsil Lee, Jaesik Choi

机构 * KAIST AI(韩国科学技术院人工智能研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments 20 pages including references and appendix; To appear in ACL 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06366 2025-06-13 q-bio.NC cs.CY cs.MA 57%

AI Agent Behavioral Science

Lin Chen, Yunke Zhang, Jie Feng, Haoye Chai, Honglin Zhang, Bingbing Fan, Yibo Ma, Shiyuan Zhang, Nian Li, Tianhui Liu, Nicholas Sukiennik, Keyu Zhao, Yu Li, Ziyi Liu, Fengli Xu, Yong Li

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02348 2025-06-13 cs.LG stat.ML 57%

Simplicity bias and optimization threshold in two-layer ReLU networks

Etienne Boursier, Nicolas Flammarion

机构 * INRIA, LMO, Université Paris-Saclay, Orsay, France(INRIA、LMO、巴黎-萨克雷大学、欧萨斯分校、法国) TML Lab, EPFL, Switzerland(TML实验室、瑞士联邦理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

Comments ICML camera ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08593 2025-06-11 cs.CL 57%

Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models

Shuzhou Yuan, Ercong Nie, Mario Tawfelis, Helmut Schmid, Hinrich Schütze, Michael Färber

机构 * ScaDS.AI and TU Dresden(ScaDS.AI 和 梵高大学) LMU Munich(慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06225 2025-06-10 cs.HC cs.AI 57%

"We need to avail ourselves of GenAI to enhance knowledge distribution": Empowering Older Adults through GenAI Literacy

Eunhye Grace Ko, Shaini Nanayakkara, Earl W. Huff

机构 * School of Information University of Texas at Austin(信息学院得克萨斯大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Journal ref CHI EA ' 2025: Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02923 2025-06-05 cs.AI stat.ML 57%

The Limits of Predicting Agents from Behaviour

Alexis Bellot, Jonathan Richens, Tom Everitt

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Journal ref ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02295 2025-06-04 cs.CL 57%

Explicit vs. Implicit: Investigating Social Bias in Large Language Models through Self-Reflection

Yachao Zhao, Bo Wang, Yan Wang, Dongming Zhao, Ruifang He, Yuexian Hou

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学) AI Lab, China Mobile Communication Group Tianjin Co., Ltd.(中国移动通信集团天津有限公司人工智能实验室)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted by ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00616 2025-06-03 cs.CY 57%

Catastrophic Liability: Managing Systemic Risks in Frontier AI Development

Aidan Kierans, Kaley Rittichier, Utku Sonsayar, Avijit Ghosh

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 10 pages, 1 figure, 1 table, in review for AIES 2025, presented at TAIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21997 2025-05-29 cs.CL 57%

Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data

Jihong Zhang, Xinya Liang, Anqi Deng, Nicole Bonge, Lin Tan, Ling Zhang, Nicole Zarrett

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05787 2025-05-29 cs.CY 57%

Mapping the Regulatory Learning Space for the EU AI Act

Dave Lewis, Marta Lasek-Markey, Delaram Golpayegani, Harshvardhan J. Pandit

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19915 2025-05-28 cs.CR cs.AI 57%

Evaluating AI cyber capabilities with crowdsourced elicitation

Artem Petrov, Dmitrii Volkov

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Updated abstract to fix a typo; no changes to the content of the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20273 2025-05-27 cs.AI 57%

Ten Principles of AI Agent Economics

Ke Yang, ChengXiang Zhai

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18657 2025-05-27 cs.AI 57%

MLLMs are Deeply Affected by Modality Bias

Xu Zheng, Chenfei Liao, Yuqian Fu, Kaiyu Lei, Yuanhuiyi Lyu, Lutao Jiang, Bin Ren, Jialei Chen, Jiawen Wang, Chengxin Li, Linfeng Zhang, Danda Pani Paudel, Xuanjing Huang, Yu-Gang Jiang, Nicu Sebe, Dacheng Tao, Luc Van Gool, Xuming Hu

机构 * HKUST(GZ)(香港科技大学(广州)) CSE, HKUST(香港科技大学计算机科学与工程系) Xi’an Jiaotong University(西安交通大学) University of Pisa, IT(比萨大学) University of Trento, IT(特伦特大学) Nagoya University(名古屋大学) China University of Mining & Technology, Beijing(中国矿业大学(北京)) Tongji University(同济大学) SPIC Energy Science and Technology Research Institute(SPIC能源科学与技术研究院) Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学) College of Computing & Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18523 2025-05-27 cs.CY 57%

Diversity and Inclusion in AI: Insights from a Survey of AI/ML Practitioners

Sidra Malik, Muneera Bano, Didar Zowghi

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13981 2025-05-27 cs.CV cs.AI 57%

On the Fairness, Diversity and Reliability of Text-to-Image Generative Models

Jordan Vice, Naveed Akhtar, Leonid Sigal, Richard Hartley, Ajmal Mian

机构 * University of Western Australia(西澳大学) University of Melbourne(墨尔本大学) University of British Columbia(不列颠哥伦比亚大学) Australian National University(澳大利亚国立大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments This research is supported by the NISDRG project #20100007, funded by the Australian Government

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17712 2025-05-26 cs.CL 57%

Understanding How Value Neurons Shape the Generation of Specified Values in LLMs

Yi Su, Jiayi Zhang, Shu Yang, Xinhai Wang, Lijie Hu, Di Wang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.19294 2025-05-23 cs.LG 57%

Investigating the Effects of Fairness Interventions Using Pointwise Representational Similarity

Camila Kolling, Till Speicher, Vedant Nanda, Mariya Toneva, Krishna P. Gummadi

机构 * MPI-SWS(马克斯·普朗克所社会科学研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13673 2025-05-21 cs.CY 57%

Comparing Apples to Oranges: A Taxonomy for Navigating the Global Landscape of AI Regulation

Sacha Alanoca, Shira Gur-Arieh, Tom Zick, Kevin Klyman

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments 24 pages, 3 figures, FAccT '25

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13466 2025-05-21 cs.AI 57%

AgentSGEN: Multi-Agent LLM in the Loop for Semantic Collaboration and GENeration of Synthetic Data

Vu Dinh Xuan, Hao Vo, David Murphy, Hoang D. Nguyen

机构 * University of Information Technology, VNU–HCM(越南胡志明市信息技术大学) University College Cork(科尔克大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12718 2025-05-20 cs.CL cs.HC 57%

Automated Bias Assessment in AI-Generated Educational Content Using CEAT Framework

Jingyang Peng, Wenyuan Shen, Jiarui Rao, Jionghao Lin

机构 * Carnegie Mellon University(卡内基梅隆大学) The University of Hong Kong(香港大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Accepted by AIED 2025: Late-Breaking Results (LBR) Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12001 2025-05-20 cs.AI cs.MA 57%

Interactional Fairness in LLM Multi-Agent Systems: An Evaluation Framework

Ruta Binkyte

机构 * Ruta Binkyte(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.07700 2025-05-15 cs.CY 57%

In Oxford Handbook on AI Governance: The Role of Workers in AI Ethics and Governance

Natalia Luka, JS Tan

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments In: Justin Bullock, Baobao Zhang, Yu-Che Chen, Johannes Himmelreich, Matthew Young, Antonin Korinek & Valerie Hudson (eds.). Oxford Handbook on AI Governance (Oxford University Press, 2022 forthcoming)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08404 2025-05-14 cs.AI 57%

Explaining Autonomous Vehicles with Intention-aware Policy Graphs

Sara Montese, Victor Gimenez-Abalos, Atia Cortés, Ulises Cortés, Sergio Alvarez-Napagao

机构 * Barcelona Supercomputing Center(巴塞罗那超级计算中心) Universitat Politècnica de Catalunya(加泰罗尼亚理工大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Accepted to Workshop EXTRAAMAS 2025 in AAMAS Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07693 2025-05-13 cs.AI 57%

Belief Injection for Epistemic Control in Linguistic State Space

Sebastian Dumbrava

机构 * Sebastian Dumbrava(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

Comments 30 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07468 2025-05-13 cs.CY 57%

Promising Topics for U.S.-China Dialogues on AI Risks and Governance

Saad Siddiqui, Lujain Ibrahim, Kristy Loke, Stephen Clare, Marianne Lu, Aris Richardson, Conor McGlynn, Jeffrey Ding

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07005 2025-05-13 cs.AI 57%

Explainable AI the Latest Advancements and New Trends

Bowen Long, Enjie Liu, Renxi Qiu, Yanqing Duan

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04393 2025-05-08 cs.CL 57%

Large Means Left: Political Bias in Large Language Models Increases with Their Number of Parameters

David Exler, Mark Schutera, Markus Reischl, Luca Rettenberger

机构 * Institute for Automation and Applied Informatics(自动化与应用信息研究所) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04291 2025-05-08 cs.CY 57%

From Incidents to Insights: Patterns of Responsibility following AI Harms

Isabel Richards, Claire Benn, Miri Zilka

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03992 2025-05-08 cs.LG 57%

Algorithmic Accountability in Small Data: Sample-Size-Induced Bias Within Classification Metrics

Jarren Briscoe, Garrett Kepler, Daryl Deford, Assefaw Gebremedhin

机构 * School of Electrical Engineering & Computer Science, Washington State University(电气工程与计算机科学学院,华盛顿州立大学) Department of Mathematics and Statistics, Washington State University(数学与统计学系,华盛顿州立大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

Comments AISTATS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏