arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2510.14053 2025-10-17 cs.AI 70%

Position: Require Frontier AI Labs To Release Small "Analog" Models

Shriyash Upadhyay, Chaithanya Bandi, Narmeen Oozeer, Philip Quirke

机构 * Frontier AI Labs(前沿人工智能实验室)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10998 2025-10-14 cs.CL cs.AI cs.CY cs.HC cs.LG 70%

ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios

Mahika Phutane, Hayoung Jung, Matthew Kim, Tanushree Mitra, Aditya Vashistha

机构 * Cornell University(康奈尔大学) Princeton University(普林斯顿大学) University of Washington(华盛顿大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 28 pages, 11 figures, 16 tables. In submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09871 2025-10-14 cs.CL 70%

CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs

Nafiseh Nikeghbal, Amir Hossein Kargaran, Jana Diesner

机构 * Technical University of Munich(慕尼黑技术大学) LMU Munich(慕尼黑大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06303 2025-10-07 cs.CY cs.AI cs.CL cs.LG 70%

On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions

Dang Nguyen, Chenhao Tan

机构 * Department of Computer Science University of Chicago(计算机科学系芝加哥大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 21 pages, 15 figures, 14 tables. Accepted as a conference paper at COLM 2025. Camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22709 2025-09-30 cs.CY cs.SY eess.SY 70%

Trust and Transparency in AI: Industry Voices on Data, Ethics, and Compliance

Louise McCormack, Diletta Huyskes, Dave Lewis, Malika Bendechache

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21075 2025-09-26 cs.CY cs.AI cs.CL cs.DC cs.HC cs.LG 70%

Communication Bias in Large Language Models: A Regulatory Perspective

Adrian Kuenzler, Stefan Schmid

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13960 2025-08-20 cs.GT cs.AI 70%

A Mechanism for Mutual Fairness in Cooperative Games with Replicable Resources -- Extended Version

Björn Filter, Ralf Möller, Özgür Lütfü Özçep

机构 * Institute for Humanities-Centered AI (CHAI), University of Hamburg, Germany(人文中心人工智能研究所(CHAI),汉堡大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

Comments This paper is the extended version of a paper accepted at the European Conference on Artificial Intelligence 2025 (ECAI 2025), providing the proof of the main theorem in the appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07623 2025-08-13 cs.CL 70%

Optimizing Class-Level Probability Reweighting Coefficients for Equitable Prompting Accuracy

Ruixi Lin, Yang You

机构 * Department of Computer Science(计算机科学系) National University of Singapore(新加坡国立大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20014 2025-07-31 cs.CR cs.AI 70%

Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation

Joydeep Chandra, Satyam Kumar Navneet

机构 * Department of CST Tsinghua University Beijing, China(计算机科学与技术系 清华大学 北京中国) Department of CSE Chandigarh University Mohali, India(计算机科学与工程系 印度昌迪加尔大学 摩哈利)

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02799 2025-07-04 cs.CL 70%

Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models

Riccardo Cantini, Nicola Gabriele, Alessio Orsino, Domenico Talia

机构 * University of Calabria, Rende, Italy(卡拉布里亚大学)

专题命中 AI治理与伦理 :safety(abstract);jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14191 2025-06-18 cs.CY 70%

The Ethics of Generative AI in Anonymous Spaces: A Case Study of 4chan's /pol/ Board

Parth Gaba, Emiliano De Cristofaro

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07282 2025-06-10 cs.CY 70%

Adultification Bias in LLMs and Text-to-Image Models

Jane Castleman, Aleksandra Korolova

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CY

Comments Accepted to the ACM Conference on Fairness, Accountability, and Transparency (FAccT '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03497 2025-06-05 cs.CY 70%

Bridging the Artificial Intelligence Governance Gap: The United States' and China's Divergent Approaches to Governing General-Purpose Artificial Intelligence

Oliver Guest, Kevin Wei

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Comments Published as a RAND commentary

Journal ref Santa Monica, CA: RAND Corporation, 2024. https://www.rand.org/pubs/perspectives/PEA3703-1.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16456 2025-06-03 cs.HC cs.AI 70%

Position: Beyond Assistance -- Reimagining LLMs as Ethical and Adaptive Co-Creators in Mental Health Care

Abeer Badawi, Md Tahmid Rahman Laskar, Jimmy Xiangji Huang, Shaina Raza, Elham Dolatabadi

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15365 2025-05-22 cs.HC cs.CL 70%

AI vs. Human Judgment of Content Moderation: LLM-as-a-Judge and Ethics-Based Response Refusals

Stefan Pasch

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00174 2025-05-07 cs.AI 70%

Real-World Gaps in AI Governance Research

Ilan Strauss, Isobel Moure, Tim O'Reilly, Sruly Rosenblat

机构 * AI Disclosures Project, Social Science Research Council Institute for Innovation and Public Purpose, University College London(AI披露项目,社会科学理事会创新与公共目的研究所,伦敦大学学院) AI Disclosures Project, Social Science Research Council(AI披露项目,社会科学理事会) O’Reilly Media(O’Reilly媒体)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

Comments Corrected a previous error: replaced 'underrepresented in Academic AI research' with the intended phrase 'underrepresented in Corporate AI research'

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01150 2025-05-05 cs.CY 70%

Methodological Foundations for AI-Driven Survey Question Generation

Ted K. Mburu, Kangxuan Rong, Campbell J. McColley, Alexandra Werth

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.CY

Comments 32 pages, 6 figures. Accepted for publication in the Journal of Engineering Education (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20086 2025-04-30 cs.CL cs.AI cs.CY cs.LG 70%

Understanding and Mitigating Risks of Generative AI in Financial Services

Sebastian Gehrmann, Claire Huang, Xian Teng, Sergei Yurovski, Iyanuoluwa Shode, Chirag S. Patel, Arjun Bhorkar, Naveen Thomas, John Doucette, David Rosenberg, Mark Dredze, David Rabinowitz

机构 * Bloomberg(彭博社) Johns Hopkins University(约翰霍普金斯大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted to FAccT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02127 2025-04-04 cs.CY 70%

Reinsuring AI: Energy, Agriculture, Finance & Medicine as Precedents for Scalable Governance of Frontier Artificial Intelligence

Nicholas Stetler

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Comments Working paper version (35 pages). Submitted to So. Ill. Law Journal; full-form citations retained for editorial review. Not peer-reviewed. Subject to revision

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22772 2025-04-01 cs.CY 70%

AI Family Integration Index (AFII): Benchmarking a New Global Readiness for AI as Family

Prashant Mahajan

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CY

Comments 28 table, 7 Figures, 78 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00265 2025-03-14 cs.AI cs.CL cs.CY cs.LG 70%

Explainable Artificial Intelligence: A Survey of Needs, Techniques, Applications, and Future Direction

Melkamu Mersha, Khang Lam, Joseph Wood, Ali AlShami, Jugal Kalita

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Journal ref 599(2024)128111

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07806 2025-03-12 cs.CL cs.AI cs.CY cs.LG 70%

Towards Large Language Models that Benefit for All: Benchmarking Group Fairness in Reward Models

Kefan Song, Jin Yao, Runnan Jiang, Rohan Chandra, Shangtong Zhang

专题命中 AI治理与伦理 :RLHF(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06411 2025-03-11 eess.SY cs.AI cs.SY 70%

Decoding the Black Box: Integrating Moral Imagination with Technical AI Governance

Krti Tallam

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09730 2025-02-18 cs.CL cs.AI cs.CY cs.HC cs.LG 70%

Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs

Huaman Sun, Jiaxin Pei, Minje Choi, David Jurgens

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07959 2025-02-04 cs.CL cs.AI cs.CY cs.LG 70%

COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act

Philipp Guldimann, Alexander Spiridonov, Robin Staab, Nikola Jovanović, Mark Vero, Velko Vechev, Anna-Maria Gueorguieva, Mislav Balunović, Nikola Konstantinov, Pavol Bielik, Petar Tsankov, Martin Vechev

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18299 2025-01-31 cs.AI 70%

Model-Free RL Agents Demonstrate System 1-Like Intentionality

Hal Ashton, Matija Franklin

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16946 2025-01-30 cs.CY 70%

Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development

Jan Kulveit, Raymond Douglas, Nora Ammann, Deger Turan, David Krueger, David Duvenaud

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Comments 19 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07780 2024-12-12 cs.CY 70%

A Taxonomy of Systemic Risks from General-Purpose AI

Risto Uuk, Carlos Ignacio Gutierrez, Daniel Guppy, Lode Lauwaert, Atoosa Kasirzadeh, Lucia Velasco, Peter Slattery, Carina Prunkl

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Comments 34 pages, 9 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16771 2024-11-22 cs.CY 70%

Navigating Governance Paradigms: A Cross-Regional Comparative Study of Generative AI Governance Processes & Principles

Jose Luna, Ivan Tan, Xiaofei Xie, Lingxiao Jiang

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.CY

Comments To appear at AIES 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12289 2024-11-18 cs.CY 70%

Catalog of General Ethical Requirements for AI Certification

Nicholas Kluge Corrêa, Julia Maria Mönig

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.CY

Comments Whitepaper - deliverable from the KI.NRW flagship-project "Zertifizierte KI"

详情

展开后加载摘要…

URL PDF HTML 收藏