arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2411.01426 2024-11-05 cs.HC cs.CY 70%

AURA: Amplifying Understanding, Resilience, and Awareness for Responsible AI Content Work

Alice Qian Zhang, Judith Amores, Mary L. Gray, Mary Czerwinski, Jina Suh

专题命中 AI治理与伦理 :safety(abstract);red teaming(abstract);分类 cs.CY

Comments To be presented at CSCW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09645 2024-10-15 cs.CY 70%

AI Model Registries: A Foundational Tool for AI Governance

Elliot McKernon, Gwyn Glasser, Deric Cheng, Gillian Hadfield

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14230 2024-09-24 cs.CY 70%

Views on AI aren't binary -- they're plural

Thorin Bristow, Luke Thorburn, Diana Acosta-Navas

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Comments 32 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17129 2024-07-25 cs.CY 70%

Mapping the individual, social, and biospheric impacts of Foundation Models

Andrés Domínguez Hernández, Shyam Krishna, Antonella Maia Perini, Michael Katell, SJ Bennett, Ann Borda, Youmna Hashem, Semeli Hadjiloizou, Sabeehah Mahomed, Smera Jayadeva, Mhairi Aitken, David Leslie

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Comments ACM Conference on Fairness, Accountability, and Transparency (FAccT '24). Association for Computing Machinery, New York, NY, USA, 776-796

Journal ref In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT '24). Association for Computing Machinery, New York, NY, USA, 776-796

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11360 2024-07-17 cs.AI 70%

Thorns and Algorithms: Navigating Generative AI Challenges Inspired by Giraffes and Acacias

Waqar Hussain

专题命中 AI治理与伦理 :safety(abstract);harmlessness(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12934 2024-06-21 cs.CR cs.AI cs.HC 70%

Current state of LLM Risks and AI Guardrails

Suriya Ganesh Ayyamperumal, Limin Ge

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

Comments Independent study, Exploring LLMs, Deploying LLMs and their Risks

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.17047 2024-05-27 cs.LG 70%

Near to Mid-term Risks and Opportunities of Open-Source Generative AI

Francisco Eiras, Aleksandar Petrov, Bertie Vidgen, Christian Schroeder de Witt, Fabio Pizzati, Katherine Elkins, Supratik Mukhopadhyay, Adel Bibi, Botos Csaba, Fabro Steibel, Fazl Barez, Genevieve Smith, Gianluca Guadagni, Jon Chun, Jordi Cabot, Joseph Marvin Imperial, Juan A. Nolazco-Flores, Lori Landay, Matthew Jackson, Paul Röttger, Philip H. S. Torr, Trevor Darrell, Yong Suk Lee, Jakob Foerster

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.LG

Comments Accepted to ICML'24 as a position paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02726 2024-05-07 cs.LG cs.SY eess.SY 70%

A Mathematical Model of the Hidden Feedback Loop Effect in Machine Learning Systems

Andrey Veprikov, Alexander Afanasiev, Anton Khritankov

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.LG

Comments 21 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16244 2024-04-30 cs.CY 70%

The Ethics of Advanced AI Assistants

Iason Gabriel, Arianna Manzini, Geoff Keeling, Lisa Anne Hendricks, Verena Rieser, Hasan Iqbal, Nenad Tomašev, Ira Ktena, Zachary Kenton, Mikel Rodriguez, Seliem El-Sayed, Sasha Brown, Canfer Akbulut, Andrew Trask, Edward Hughes, A. Stevie Bergman, Renee Shelby, Nahema Marchal, Conor Griffin, Juan Mateos-Garcia, Laura Weidinger, Winnie Street, Benjamin Lange, Alex Ingerman, Alison Lentz, Reed Enger, Andrew Barakat, Victoria Krakovna, John Oliver Siy, Zeb Kurth-Nelson, Amanda McCroskery, Vijay Bolina, Harry Law, Murray Shanahan, Lize Alberts, Borja Balle, Sarah de Haas, Yetunde Ibitoye, Allan Dafoe, Beth Goldberg, Sébastien Krier, Alexander Reese, Sims Witherspoon, Will Hawkins, Maribeth Rauh, Don Wallace, Matija Franklin, Josh A. Goldstein, Joel Lehman, Michael Klenk, Shannon Vallor, Courtney Biles, Meredith Ringel Morris, Helen King, Blaise Agüera y Arcas, William Isaac, James Manyika

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14683 2024-03-25 cs.CY cs.AI cs.CL cs.LG 70%

A Moral Imperative: The Need for Continual Superalignment of Large Language Models

Gokul Puthumanaillam, Manav Vora, Pranay Thangeda, Melkior Ornik

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05553 2023-10-10 cs.CL 70%

Regulation and NLP (RegNLP): Taming Large Language Models

Catalina Goanta, Nikolaos Aletras, Ilias Chalkidis, Sofia Ranchordas, Gerasimos Spanakis

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL

Comments 9 pages, long paper at EMNLP 2023 proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.07120 2023-09-14 cs.CL cs.AI cs.CV cs.CY cs.LG 70%

Sight Beyond Text: Multi-Modal Training Enhances LLMs in Truthfulness and Ethics

Haoqin Tu, Bingchen Zhao, Chen Wei, Cihang Xie

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.11525 2023-07-24 cs.AI 70%

Model Reporting for Certifiable AI: A Proposal from Merging EU Regulation into AI Development

Danilo Brajovic, Niclas Renner, Vincent Philipp Goebels, Philipp Wagner, Benjamin Fresz, Martin Biller, Mara Klaeb, Janika Kutz, Jens Neuhuettler, Marco F. Huber

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI

Comments 54 pages, 1 figure, to be submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.04699 2023-07-12 cs.CY 70%

International Institutions for Advanced AI

Lewis Ho, Joslyn Barnhart, Robert Trager, Yoshua Bengio, Miles Brundage, Allison Carnegie, Rumman Chowdhury, Allan Dafoe, Gillian Hadfield, Margaret Levi, Duncan Snidal

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Comments 19 pages, 2 figures, fixed rendering issues

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.03279 2023-06-14 cs.LG cs.AI cs.CL cs.CY 70%

Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Jonathan Ng, Hanlin Zhang, Scott Emmons, Dan Hendrycks

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments ICML 2023 Oral (camera-ready); 31 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.06749 2023-06-13 cs.CY 70%

Implementing AI Ethics: Making Sense of the Ethical Requirements

Mamia Agbese, Rahul Mohanani, Arif Ali Khan, Pekka Abrahamsson

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09941 2023-06-05 cs.CL cs.AI cs.CY cs.LG 70%

"I'm fully who I am": Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation

Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, Rahul Gupta

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Journal ref 2023 ACM Conference on Fairness, Accountability, and Transparency

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.10831 2023-05-03 cs.CY cs.HC 70%

Bridging Deliberative Democracy and Deployment of Societal-Scale Technology

Ned Cooper

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

Comments 3 pages

Journal ref CHI 2023 Workshop on Designing Technology and Policy Simultaneously

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00868 2022-07-05 cs.AI cs.CL cs.CY cs.LG 70%

The Linguistic Blind Spot of Value-Aligned Agency, Natural and Artificial

Travis LaCroix

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 49 pages; 2 Figures; 1 Table; -- Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.07369 2022-05-17 cs.AI cs.MA math.DS nlin.AO 70%

Understanding Emergent Behaviours in Multi-Agent Systems with Evolutionary Game Theory

The Anh Han

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.00672 2021-10-05 cs.CY cs.AI cs.CL cs.LG 70%

Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models

Robert Wolfe, Aylin Caliskan

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 15 pages, 3 figures, 8 tables

Journal ref Empirical Methods in Natural Language Processing 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.02117 2021-05-06 cs.CY 70%

Ethics and Governance of Artificial Intelligence: Evidence from a Survey of Machine Learning Researchers

Baobao Zhang, Markus Anderljung, Lauren Kahn, Noemi Dreksler, Michael C. Horowitz, Allan Dafoe

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2002.05672 2020-11-11 cs.CY 70%

Steps Towards Value-Aligned Systems

Osonde A. Osoba, Benjamin Boudreaux, Douglas Yeung

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CY

Comments Original version appeared in Proceedings of the 2020 AAAI ACM Conference on AI, Ethics, and Society (AIES '20), February 7-8, 2020, New York, NY, USA. 5 pages, 2 figures. Corrected some typos in this version

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.02817 2017-06-06 cs.AI 70%

Unethical Research: How to Create a Malevolent Artificial Intelligence

Federico Pistono, Roman V. Yampolskiy

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI

Journal ref In proceedings of Ethics for Artificial Intelligence Workshop (AI-Ethics-2016). Pages 1-7. New York, NY. July 9 -- 15, 2016

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16960 2023-10-31 cs.CL cs.AI cs.CY cs.HC 69%

Training Socially Aligned Language Models on Simulated Social Interactions

Ruibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang, Denny Zhou, Andrew M. Dai, Diyi Yang, Soroush Vosoughi

专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CL、cs.AI、cs.CY

Comments Code, data, and models can be downloaded via https://github.com/agi-templar/Stable-Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11924 2026-08-24 cs.CL cs.AI cs.LG 版本更新 67%

Explaining Intrinsic Moral Self-Correction with Mechanistic Interpretability

大语言模型中的内在自我修正:通过机制可解释性实现可解释的提示

Yu-Ting Lee, Fu-Chieh Chang, Yu-En Shu, Hui-Ying Shih, Pei-Yuan Wu

机构 * Graduate Institute of Communication Engineering, National Taiwan University, Taipei, Taiwan(通讯工程研究院,国立台湾大学) MediaTek Research, Taipei, Taiwan(联发科研究,台北,台湾) Department of Electrical Engineering, National Taiwan University, Taipei, Taiwan(电气工程系,国立台湾大学) Department of Electrical Engineering, National Tsing Hua University, Hsinchu, Taiwan(电气工程系,国立清华大学) AI Research Center (AINTU), National Taiwan University, Taipei, Taiwan(人工智能研究中心(AINTU),国立台湾大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本研究通过机制可解释性揭示大语言模型中内在自我修正的机制,证明提示引导隐藏表示偏移是其核心驱动因素。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11251 2026-08-13 cs.CY cs.AI cs.LG 新提交 67%

Variable Selection in the Context of AI Fairness

AI公平性语境下的变量选择

Ivan Luciano Danesi, Chiara Frigerio, Fabio Maccaferri, Giorgio Alessandro Motta, Pietro Zecca

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文针对AI公平性语境下的变量选择问题,提出将数学方法与伦理、社会意识融合的跨学科方案,主张保留所有潜在相关变量以减少隐含偏差,助力AI系统符合欧盟AI法案要求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06112 2026-08-07 cs.AI cs.CL cs.LG cs.MA 新提交 67%

From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems

从孤立算法到合规优先的智能体平台:面向医院AI系统的多层架构

Manideep Dhar, Ritwik Singh, Sharat Chandra Kumar Manikonda

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 针对医院AI部署孤立、规模化难的问题,提出含智能体编排、合规策略、隐私保护数据层的多层架构,通过原型验证可减少任务耗时与文档工作量,为医院AI平台建设提供实用蓝图。

Comments Peer-reviewed published article

Journal ref IJISRT, 11-2026(5), IJISRT26MAY1651

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28934 2026-08-03 cs.CL cs.AI cs.CY 新提交 67%

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

FairFund-Bench:评估大语言模型资源分配中的分配偏差

Martin Lukk

机构 * University of Toronto(多伦多大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 FairFund-Bench基准通过调整审计格式等特征,发现LLM资源分配的偏差受审计类型影响,且因果框架效应强于人口统计效应,可用于评估LLM的分配偏差。

Comments 19 pages, 7 figures. Code and data: https://github.com/martinlukk/fairfund-bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08789 2026-07-29 cs.CR 版本更新 67%

Never compromise with vulnerabilities: a comprehensive survey on AI governance

绝不向漏洞妥协:人工智能治理综合调查

Yuchu Jiang, Jian Zhao, Yuchen Yuan, Tianle Zhang, Yao Huang, Yanghao Zhang, Yan Wang, Yanshu Li, Xizhong Guo, Yusheng Zhao, Huilin Zhou, Jun Zhang, Zhi Zhang, Xiaojian Lin, Yixiu Zou, Haoxuan Ma, Yuhu Shang, Yuzhi Hu, Keshu Cai, Ruochen Zhang, Boyuan Chen, Yilan Gao, Ziheng Jiao, Yi Qin, Shuangjun Du, Xiao Tong, Zhekun Liu, Yu Chen, Xuankun Rong, Rui Wang, Yejie Zheng, Zhaoxin Fan, Murat Sensoy, Hongyuan Zhang, Pan Zhou, Lei Jin, Hao Zhao, Xu Yang, Jiaojiao Zhao, Jianshu Li, Joey Tianyi Zhou, Zhi-Qi Cheng, Longtao Huang, Zhiyi Liu, Zheng Zhu, Jianan Li, Gang Wang, Qi Li, Xu-Yao Zhang, Yaodong Yang, Mang Ye, Wenqi Ren, Zhaofeng He, Hang Su, Rongrong Ni, Liping Jing, Xingxing Wei, Junliang Xing, Massimo Alioto, Shengmei Shen, Petia Radeva, Dacheng Tao, Ya-Qin Zhang, Shuicheng Yan, Xuelong Li

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract)

AI总结 针对人工智能发展带来的技术漏洞及社会风险,提出整合技术与社会维度的框架,围绕三个支柱构建,通过系统回顾识别核心挑战,给出综合研究议程,为开发可靠且符合伦理的人工智能系统提供指导。

Comments 26 pages, 3 figures, accepted by SCIS2026

详情

展开后加载摘要…

URL PDF HTML 收藏