arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2510.09634 2025-10-14 cs.CY cs.AI 62%

Responsible AI Adoption in the Public Sector: A Data-Centric Taxonomy of AI Adoption Challenges

Anastasija Nikiforova, Martin Lnenicka, Ulf Melin, David Valle-Cruz, Asif Gill, Cesar Casiano Flores, Emyana Sirait, Mariusz Luterek, Richard Michael Dreyling, Barbora Tesarova

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07613 2025-10-10 cs.CL cs.AI 62%

Vocabulary embeddings organize linguistic structure early in language model training

Isabel Papadimitriou, Jacob Prince

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02563 2025-10-08 cs.LG cs.CL 62%

DynaGuard: A Dynamic Guardian Model With User-Defined Policies

Monte Hoover, Vatsal Baherwani, Neel Jain, Khalid Saifullah, Joseph Vincent, Chirag Jain, Melissa Kazemi Rad, C. Bayan Bruss, Ashwinee Panda, Tom Goldstein

机构 * University of Maryland(马里兰大学) Capital One

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.LG

Comments 22 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10659 2025-10-07 cs.SI cs.AI cs.CL cs.MA 62%

Network Formation and Dynamics Among Multi-LLMs

Marios Papachristou, Yuan Yuan

机构 * Department of Information Systems, W.P. Carey School of Business, Arizona State University, Tempe, AZ, USA(亚利桑那州立大学信息系统系,W.P. Carey商学院,Tempe分校) Department of Computer Science, Cornell University, Ithaca, NY, USA(康奈尔大学计算机科学系) Graduate School of Management, University of California Davis, Davis, CA, USA(加州大学戴维斯分校管理研究生院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at PNAS Nexus

Journal ref PNAS Nexus 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03368 2025-10-07 cs.CY cs.AI 62%

An Adaptive Responsible AI Governance Framework for Decentralized Organizations

Kiana Jafari Meimandi, Anka Reuel, Gabriela Aranguiz-Dias, Hatim Rahama, Ala-Eddine Ayadi, Xavier Boullier, Jérémy Verdo, Louis Montanie, Mykel Kochenderfer

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03004 2025-10-06 cs.LG cs.AI 62%

BrainIB++: Leveraging Graph Neural Networks and Information Bottleneck for Functional Brain Biomarkers in Schizophrenia

Tianzheng Hu, Qiang Li, Shu Liu, Vince D. Calhoun, Guido van Wingen, Shujian Yu

机构 * Vrije University Amsterdam(荷兰阿姆斯特丹自由大学) Tri-institutional Center for Translational Research in Neuroimaging(转化神经影像研究联合中心) Emory University(埃默里大学) Key Laboratory of Genetic Evolution and Animal Models(遗传进化与动物模型重点实验室) Kunming Institute of Zoology(昆明动物研究所) Chinese Academy of Sciences Kunming(中国科学院昆明分院) Department of Psychiatry, Amsterdam UMC, University of Amsterdam(阿姆斯特丹大学精神病科) Department of Physics and Technology, UiT The Arctic University of Norway(北极大学挪威理工学院物理与技术系)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments This manuscript has been accepted by Biomedical Signal Processing and Control and the code is available at https://github.com/TianzhengHU/BrainIB_coding/tree/main/BrainIB_GIB

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02978 2025-10-06 cs.CY cs.AI cs.HC 62%

AI Generated Child Sexual Abuse Material -- What's the Harm?

Caoilte Ó Ciardha, John Buckley, Rebecca S. Portnoff

机构 * Senior Research Fellow, University of Kent, UK(肯特大学高级研究员) Digital Child Safety Expert(数字儿童安全专家) Vice President of Data Science, Thorn(数据科学副总裁,Thorn)

专题命中 AI治理与伦理 :harmlessness(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22774 2025-10-06 cs.AI cs.CY 62%

Bridging Ethical Principles and Algorithmic Methods: An Alternative Approach for Assessing Trustworthiness in AI Systems

Michael Papademas, Xenia Ziouvelou, Antonis Troumpoukis, Vangelis Karkaletsis

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26584 2025-10-01 cs.AI cs.IR cs.LG cs.SE 62%

Fairness Testing in Retrieval-Augmented Generation: How Small Perturbations Reveal Bias in Small Language Models

Matheus Vinicius da Silva de Oliveira, Jonathan de Andrade Silva, Awdren de Lima Fontao

机构 * Faculty of Computing - Federal University of Mato Grosso do Sul(计算机学院 - 短暂戈亚那联邦大学)

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07931 2025-10-01 cs.CY cs.AI 62%

Educating a Responsible AI Workforce: Piloting a Curricular Module on AI Policy in a Graduate Machine Learning Course

James Weichert, Hoda Eldardiry

机构 * Department of Computer Science Virginia Tech(计算机科学系弗吉尼亚理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Accepted at 2025 ASEE Annual Conference & Exposition

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16355 2025-09-29 cs.LG cs.AI 62%

How Strategic Agents Respond: Comparing Analytical Models with LLM-Generated Responses in Strategic Classification

Tian Xie, Pavan Rauch, Xueru Zhang

机构 * The Ohio State University(俄亥俄州立大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Add GPT 5 experiments

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.17805 2025-09-29 cs.CY cs.AI 62%

Biospheric AI

Marcin Korecki

机构 * TU Delft(代尔夫特理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10653 2025-09-16 cs.CY cs.AI 62%

SCOR: A Framework for Responsible AI Innovation in Digital Ecosystems

Mohammad Saleh Torkestani, Taha Mansouri

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Proceeding of The British Academy of Management Conference 2025, University of Kent, UK

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10289 2025-09-15 cs.CY cs.AI 62%

We Need a New Ethics for a World of AI Agents

Iason Gabriel, Geoff Keeling, Arianna Manzini, James Evans

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 6 pages, no figures

Journal ref Nature, 644 (8075), 2025, 38-40

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02773 2025-09-15 cs.CY cs.AI econ.GN q-fin.EC 62%

Web3 x AI Agents: Landscape, Integrations, and Foundational Challenges

Yiming Shen, Jiashuo Zhang, Zhenzhe Shao, Wenxuan Luo, Yanlin Wang, Ting Chen, Zibin Zheng, Jiachi Chen

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08829 2025-09-12 cs.CY cs.AI cs.IR 62%

PerFairX: Is There a Balance Between Fairness and Personality in Large Language Model Recommendations?

Chandan Kumar Sah

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 10 pages, 5 figures. Accepted to the Workshop on Multimodal Continual Learning (MCL) at ICCV 2025. @2025 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), ICCV's 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22591 2025-09-09 cs.LG cs.AI stat.ME 62%

FACEGroup: Feasible and Actionable Counterfactual Explanations for Group Fairness

Christos Fragkathoulas, Vasiliki Papanikou, Evaggelia Pitoura, Evimaria Terzi

机构 * University of Ioannina(伊奥安纳大学) Archimedes, Athena Research Center(阿基米德研究所) Boston University(波士顿大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments ECML PKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05627 2025-09-09 cs.CY cs.LG stat.ML 62%

Audits Under Resource, Data, and Access Constraints: Scaling Laws For Less Discriminatory Alternatives

Sarah H. Cen, Salil Goyal, Zaynah Javed, Ananya Karthik, Percy Liang, Daniel E. Ho

机构 * Stanford University(斯坦福大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

Comments 34 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08193 2025-09-05 cs.CY cs.AI 62%

Street-Level AI: Are Large Language Models Ready for Real-World Judgments?

Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler, Sanmay Das

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments This work has been accepted for publication as a full paper at the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01576 2025-09-03 cs.AI cs.CY cs.SY eess.SY 62%

Structured AI Decision-Making in Disaster Management

Julian Gerald Dcruz, Argyrios Zolotas, Niall Ross Greenwood, Miguel Arana-Catania

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 40 pages, 14 figures, 16 tables. To be published in Nature Scientific Reports

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21788 2025-09-01 cs.CL cs.AI cs.IR 62%

Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval

Inés Altemir Marinas, Anastasiia Kucherenko, Andrei Kucharavy

机构 * École Polytechnique Fédérale de Lausanne(瑞士联邦理工学院) Institute of Entrepreneurship and Management, HES-SO Valais-Wallis(创业与管理研究所) Institute of Informatics, HES-SO Valais-Wallis(信息研究所)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20015 2025-08-28 cs.LG cs.AI 62%

Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment

Julian Arnold, Niels Lörch

机构 * Department of Physics University of Basel(物理系 巴塞尔大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments 11+25 pages, 4+11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16762 2025-08-26 cs.CL cs.CY 62%

Toward Socially Aware Vision-Language Models: Evaluating Cultural Competence Through Multimodal Story Generation

Arka Mukherjee, Shreya Ghosh

机构 * KIIT Deemed University(KIIT大学) Indian Institute of Technology (IIT) Bhubaneswar(印度理工学院(Bhubaneswar分校))

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

Comments Accepted at ASI @ ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14116 2025-08-21 cs.CY cs.AI 62%

Enriching Moral Perspectives on AI: Concepts of Trust amongst Africans

Lameck Mbangula Amugongo, Nicola J Bidwell, Joseph Mwatukange

机构 * Namibia University of Science \& Technology 13 Jackson Kaujeua Windhoek Namibia 9000 Rhodes University Makhanda South Africa International University of Management Namibia Charles Darwin University Australia Namibia University of Science \& Technology Rhodes University International University of Management Charles Darwin University

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08804 2025-08-13 cs.LG cs.AI 62%

TechOps: Technical Documentation Templates for the AI Act

Laura Lucaj, Alex Loosley, Hakan Jonsson, Urs Gasser, Patrick van der Smagt

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08544 2025-08-13 cs.CY cs.AI 62%

AI Agents and the Law

Mark O. Riedl, Deven R. Desai

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 2025 AAAI Conference on AI, Ethics, and Society

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05269 2025-08-13 cs.LG cs.AI q-bio.QM 62%

Chemist-aligned retrosynthesis by ensembling diverse inductive bias models

Krzysztof Maziarz, Guoqing Liu, Hubert Misztela, Austin Tripp, Junren Li, Aleksei Kornev, Piotr Gaiński, Holger Hoefling, Mike Fortunato, Rishi Gupta, Marwin Segler

机构 * Microsoft Research AI for Science(微软研究院人工智能与科学研究中心) Novartis Biomedical Research(诺华生物医学研究) University of Cambridge(剑桥大学) Jagiellonian University(雅盖隆大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08333 2025-08-13 cs.CY cs.AI 62%

Normative Moral Pluralism for AI: A Framework for Deliberation in Complex Moral Contexts

David-Doron Yaacov

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Conference version: AIES 2025 (non-archival track), 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07673 2025-08-12 cs.AI cs.LG 62%

Ethics2vec: aligning automatic agents and human preferences

Gianluca Bontempi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07111 2025-08-12 cs.CL cs.AI 62%

Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution

Falaah Arif Khan, Nivedha Sivakumar, Yinong Oliver Wang, Katherine Metcalf, Cezanne Camacho, Barry-John Theobald, Luca Zappella, Nicholas Apostoloff

机构 * Apple(苹果公司)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏