arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2310.18333 2023-12-18 cs.CL cs.AI 62%

She had Cobalt Blue Eyes: Prompt Testing to Create Aligned and Sustainable Language Models

Veronica Chatrath, Oluwanifemi Bamgbose, Shaina Raza

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted as Oral at the AAAI 2nd Workshop on Sustainable AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.15936 2023-12-01 cs.CY cs.LG 62%

Towards Responsible Governance of Biological Design Tools

Richard Moulange, Max Langenkamp, Tessa Alexanian, Samuel Curtis, Morgan Livingston

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

Comments 10 pages + references, 1 figure, accepted at NeurIPS 2023 Workshop on Regulatable ML as oral presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.14684 2023-11-28 cs.CY cs.AI 62%

The risks of risk-based AI regulation: taking liability seriously

Martin Kretschmer, Tobias Kretschmer, Alexander Peukert, Christian Peukert

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07141 2023-11-17 cs.LG cs.CY 62%

SABAF: Removing Strong Attribute Bias from Neural Networks with Adversarial Filtering

Jiazhi Li, Mahyar Khayatkhoei, Jiageng Zhu, Hanchen Xie, Mohamed E. Hussein, Wael AbdAlmageed

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG

Comments 35 pages, 18 figures, 32 tables. This work is an extended version of our paper (arXiv:2310.04955). Code will be released at https://github.com/jiazhi412/strong_attribute_bias

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02294 2023-11-07 cs.CL cs.CY 62%

LLMs grasp morality in concept

Mark Pock, Andre Ye, Jared Moore

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

Comments Presented at NeurIPS 2023 Moral Pyschology and Moral Philosophy workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08901 2023-10-16 cs.MA cs.AI cs.CL 62%

Welfare Diplomacy: Benchmarking Language Model Cooperation

Gabriel Mukobi, Hannah Erlebach, Niklas Lauffer, Lewis Hammond, Alan Chan, Jesse Clifton

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.02796 2023-10-12 cs.DB cs.CL cs.LG 62%

VerifAI: Verified Generative AI

Nan Tang, Chenyu Yang, Ju Fan, Lei Cao, Yuyu Luo, Alon Halevy

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.LG

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05840 2023-10-10 cs.LG cs.AI 62%

Predicting Accident Severity: An Analysis Of Factors Affecting Accident Severity Using Random Forest Model

Adekunle Adefabi, Somtobe Olisah, Callistus Obunadike, Oluwatosin Oyetubo, Esther Taiwo, Edward Tella

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments 15 pages

Journal ref International Journal on Cybernetics & Informatics (IJCI) Vol.12, No.6, December 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.04963 2023-09-29 cs.AI cs.CY cs.SE 62%

Responsible AI Pattern Catalogue: A Collection of Best Practices for AI Governance and Engineering

Qinghua Lu, Liming Zhu, Xiwei Xu, Jon Whittle, Didar Zowghi, Aurelie Jacquet

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12356 2023-09-25 cs.CY cs.AI 62%

A Critical Examination of the Ethics of AI-Mediated Peer Review

Laurie A. Schintler, Connie L. McNeely, James Witte

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 21 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.10318 2023-09-20 cs.AI cs.CY 62%

Who to Trust, How and Why: Untangling AI Ethics Principles, Trustworthiness and Trust

Andreas Duenser, David M. Douglas

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 7 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.09979 2023-08-22 cs.CY cs.AI 62%

Artificial Intelligence across Europe: A Study on Awareness, Attitude and Trust

Teresa Scantamburlo, Atia Cortés, Francesca Foffano, Cristian Barrué, Veronica Distefano, Long Pham, Alessandro Fabris

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03198 2023-07-14 cs.CY cs.AI 62%

A multilevel framework for AI governance

Hyesun Choung, Prabu David, John S. Seberger

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments This paper has been accepted for publication and is forthcoming in The Global and Digital Governance Handbook. Cite as: Choung, H., David, P., & Seberger, J.S. (2023). A multilevel framework for AI governance. The Global and Digital Governance Handbook. Routledge, Taylor & Francis Group

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.01800 2023-06-06 cs.CY cs.AI 62%

The ethical ambiguity of AI data enrichment: Measuring gaps in research ethics norms and practices

Will Hawkins, Brent Mittelstadt

专题命中 AI治理与伦理 :RLHF(abstract);分类 cs.AI、cs.CY

Comments 10 pages

Journal ref 2023 ACM Conference on Fairness, Accountability, and Transparency

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.08486 2023-05-30 cs.CV cs.AI cs.LG 62%

Scalar Invariant Networks with Zero Bias

Chuqin Geng, Xiaojie Xu, Haolin Ye, Xujie Si

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 22 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.17137 2023-05-30 cs.AI cs.LG 62%

Integrating Generative Artificial Intelligence in Intelligent Vehicle Systems

Lukas Stappen, Jeremy Dillmann, Serena Striegel, Hans-Jörg Vögel, Nicolas Flores-Herr, Björn W. Schuller

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.11215 2023-04-25 cs.CY cs.AI 62%

ChatGPT: More than a Weapon of Mass Deception, Ethical challenges and responses from the Human-Centered Artificial Intelligence (HCAI) perspective

Alejo Jose G. Sison, Marco Tulio Daza, Roberto Gozalo-Brizuela, Eduardo C. Garrido-Merchán

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.03843 2023-03-07 cs.LG cs.CY 62%

Counterfactual Fairness Is Basically Demographic Parity

Lucas Rosenblatt, R. Teal Witter

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.05442 2023-02-13 cs.CV cs.AI cs.LG 62%

Scaling Vision Transformers to 22 Billion Parameters

Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, Rodolphe Jenatton, Lucas Beyer, Michael Tschannen, Anurag Arnab, Xiao Wang, Carlos Riquelme, Matthias Minderer, Joan Puigcerver, Utku Evci, Manoj Kumar, Sjoerd van Steenkiste, Gamaleldin F. Elsayed, Aravindh Mahendran, Fisher Yu, Avital Oliver, Fantine Huot, Jasmijn Bastings, Mark Patrick Collier, Alexey Gritsenko, Vighnesh Birodkar, Cristina Vasconcelos, Yi Tay, Thomas Mensink, Alexander Kolesnikov, Filip Pavetić, Dustin Tran, Thomas Kipf, Mario Lučić, Xiaohua Zhai, Daniel Keysers, Jeremiah Harmsen, Neil Houlsby

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.02323 2023-02-07 cs.LG cs.AI stat.ML 62%

Improving Fair Training under Correlation Shifts

Yuji Roh, Kangwook Lee, Steven Euijong Whang, Changho Suh

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.00479 2023-01-18 cs.CR cs.AI cs.CL stat.ML 62%

The Design Principle of Blockchain: An Initiative for the SoK of SoKs

Luyao Sunshine Zhang

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.08704 2022-12-01 cs.LG cs.CY 62%

Accurate Fairness: Improving Individual Fairness without Trading Accuracy

Xuran Li, Peng Wu, Jing Su

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03274 2022-11-24 cs.LG cs.AI 62%

TCNL: Transparent and Controllable Network Learning Via Embedding Human-Guided Concepts

Zhihao Wang, Chuang Zhu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.15289 2022-10-28 cs.CY cs.AI 62%

On the Efficiency of Ethics as a Governing Tool for Artificial Intelligence

Nicholas Kluge Corrêa, Nythamar De Oliveira, Diogo Massmann

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.14975 2022-10-28 cs.CL cs.LG 62%

MABEL: Attenuating Gender Bias using Textual Entailment Data

Jacqueline He, Mengzhou Xia, Christiane Fellbaum, Danqi Chen

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted to EMNLP 2022. Code and models are publicly available at https://github.com/princeton-nlp/mabel

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.02667 2022-10-07 cs.AI cs.CY 62%

A Human Rights-Based Approach to Responsible AI

Vinodkumar Prabhakaran, Margaret Mitchell, Timnit Gebru, Iason Gabriel

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Presented as a (non-archival) poster at the 2022 ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization or (EAAMO '22)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.11770 2022-09-27 cs.HC cs.AI cs.LG 62%

Toward Smart Doors: A Position Paper

Luigi Capogrosso, Geri Skenderi, Federico Girella, Franco Fummi, Marco Cristani

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments 2nd International Workshop on Industrial Machine Learning @ ICPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.12645 2022-08-29 cs.CY cs.AI 62%

The Brussels Effect and Artificial Intelligence: How EU regulation will impact the global AI market

Charlotte Siegmann, Markus Anderljung

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.07635 2022-08-19 cs.AI cs.CY 62%

AI Ethics Issues in Real World: Evidence from AI Incident Database

Mengyi Wei, Zhixuan Zhou

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 56th Hawaii International Conference on System Sciences (HICSS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.03721 2022-06-14 cs.CY cs.AI 62%

Demystifying the Draft EU Artificial Intelligence Act

Michael Veale, Frederik Zuiderveen Borgesius

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 16 pages, 1 table

Journal ref Computer Law Review International (2021), 22(4) 97-112

详情

展开后加载摘要…

URL PDF HTML 收藏