arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2401.12273 2024-07-11 cs.CR cs.AI cs.CL 62%

The Ethics of Interaction: Mitigating Security Threats in LLMs

Ashutosh Kumar, Shiv Vignesh Murthy, Sagarika Singh, Swathy Ragupathy

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.06235 2024-07-10 cs.CY cs.AI 62%

Auditing of AI: Legal, Ethical and Technical Approaches

Jakob Mokander

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Journal ref DISO 2, 49 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19497 2024-07-01 cs.CL cs.AI 62%

Inclusivity in Large Language Models: Personality Traits and Gender Bias in Scientific Abstracts

Naseela Pervez, Alexander J. Titus

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09779 2024-06-17 cs.AI cs.CL cs.CV 62%

OSPC: Detecting Harmful Memes with Large Language Model as a Catalyst

Jingtao Cao, Zheng Zhang, Hongru Wang, Bin Liang, Hao Wang, Kam-Fai Wong

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06435 2024-06-11 cs.CL cs.AI 62%

Language Models are Alignable Decision-Makers: Dataset and Application to the Medical Triage Domain

Brian Hu, Bill Ray, Alice Leung, Amy Summerville, David Joy, Christopher Funk, Arslan Basharat

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments 15 pages total (including appendix), NAACL 2024 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10819 2024-06-11 cs.LG cs.AI stat.ML 62%

Auditing and Generating Synthetic Data with Controllable Trust Trade-offs

Brian Belgodere, Pierre Dognin, Adam Ivankay, Igor Melnyk, Youssef Mroueh, Aleksandra Mojsilovic, Jiri Navratil, Apoorva Nitsure, Inkit Padhi, Mattia Rigotti, Jerret Ross, Yair Schiff, Radhika Vedpathak, Richard A. Young

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments submitted

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04671 2024-06-10 cs.CY cs.AI 62%

The Reasonable Person Standard for AI

Sunayana Rane

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03299 2024-06-06 cs.AI cs.CL 62%

The Good, the Bad, and the Hulk-like GPT: Analyzing Emotional Decisions of Large Language Models in Cooperation and Bargaining Games

Mikhail Mozikov, Nikita Severin, Valeria Bodishtianu, Maria Glushanina, Mikhail Baklashkin, Andrey V. Savchenko, Ilya Makarov

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13041 2024-06-06 cs.CL cs.AI 62%

Assessing Political Bias in Large Language Models

Luca Rettenberger, Markus Reischl, Mark Schutera

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01264 2024-05-24 cs.AI cs.CL 62%

Exploring the psychology of LLMs' Moral and Legal Reasoning

Guilherme F. C. F. Almeida, José Luiz Nunes, Neele Engelmann, Alex Wiegmann, Marcelo de Araújo

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref Exploring the psychology of LLMs' moral and legal reasoning. Artificial Intelligence, Volume 224, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.11273 2024-05-21 cs.AI cs.CL cs.CV cs.MM 62%

Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts

Yunxin Li, Shenyuan Jiang, Baotian Hu, Longyue Wang, Wanqi Zhong, Wenhan Luo, Lin Ma, Min Zhang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 22 pages, 13 figures. Project Website: https://uni-moe.github.io/. Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07076 2024-05-15 cs.CL cs.AI 62%

Integrating Emotional and Linguistic Models for Ethical Compliance in Large Language Models

Edward Y. Chang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments 29 pages, 10 tables, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.14660 2024-04-24 cs.CY cs.AI 62%

AI Procurement Checklists: Revisiting Implementation in the Age of AI Governance

Tom Zick, Mason Kortz, David Eaves, Finale Doshi-Velez

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.11271 2024-04-02 cs.CL cs.CY cs.HC 62%

MONAL: Model Autophagy Analysis for Modeling Human-AI Interactions

Shu Yang, Muhammad Asif Ali, Lu Yu, Lijie Hu, Di Wang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.06153 2024-03-29 cs.LG cs.AI cs.HC 62%

Open Datasheets: Machine-readable Documentation for Open Datasets and Responsible AI Assessments

Anthony Cintron Roman, Jennifer Wortman Vaughan, Valerie See, Steph Ballard, Jehu Torres, Caleb Robinson, Juan M. Lavista Ferres

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17368 2024-03-27 cs.CL cs.AI 62%

ChatGPT Rates Natural Language Explanation Quality Like Humans: But on Which Scales?

Fan Huang, Haewoon Kwak, Kunwoo Park, Jisun An

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accpeted by LREC-COLING 2024 main conference, long paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.15601 2024-03-26 cs.CY cs.AI 62%

From Guidelines to Governance: A Study of AI Policies in Education

Aashish Ghimire, John Edwards

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.14469 2024-03-22 cs.CL cs.AI 62%

ChatGPT Alternative Solutions: Large Language Models Survey

Hanieh Alipour, Nick Pendar, Kohinoor Roy

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Journal ref David C. Wyld et al. (Eds): NBIoT, MLCL, NMCO, ARIN, CSITA, ISPR, NATAP-2024. pp. 153-173, 2024. CS & IT - CSCP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13840 2024-03-22 cs.CL cs.AI cs.SI 62%

Whose Side Are You On? Investigating the Political Stance of Large Language Models

Pagnarasmey Pit, Xingjun Ma, Mike Conway, Qingyu Chen, James Bailey, Henry Pit, Putrasmey Keo, Watey Diep, Yu-Gang Jiang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.11402 2024-03-19 cs.CY cs.AI 62%

Embracing the Generative AI Revolution: Advancing Tertiary Education in Cybersecurity with GPT

Raza Nowrozy, David Jam

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09510 2024-03-15 cs.AI cs.CY cs.GT cs.MA math.DS 62%

Trust AI Regulation? Discerning users are vital to build trust and effective AI regulation

Zainab Alalawi, Paolo Bova, Theodor Cimpeanu, Alessandro Di Stefano, Manh Hong Duong, Elias Fernandez Domingos, The Anh Han, Marcus Krellner, Bianca Ogbo, Simon T. Powers, Filippo Zimmaro

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.00064 2024-03-01 cs.CY cs.AI 62%

Ethical Framework for Harnessing the Power of AI in Healthcare and Beyond

Sidra Nasir, Rizwan Ahmed Khan, Samita Bai

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Journal ref IEEE Access 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13379 2024-02-22 cs.LG cs.CY 62%

Referee-Meta-Learning for Fast Adaptation of Locational Fairness

Weiye Chen, Yiqun Xie, Xiaowei Jia, Erhu He, Han Bao, Bang An, Xun Zhou

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04489 2024-02-08 cs.LG cs.CR cs.CY stat.ME 62%

De-amplifying Bias from Differential Privacy in Language Model Fine-tuning

Sanjari Srivastava, Piotr Mardziel, Zhikhun Zhang, Archana Ahlawat, Anupam Datta, John C Mitchell

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.07353 2024-02-06 cs.SE cs.AI cs.LG 62%

Towards Engineering Fair and Equitable Software Systems for Managing Low-Altitude Airspace Authorizations

Usman Gohar, Michael C. Hunter, Agnieszka Marczak-Czajka, Robyn R. Lutz, Myra B. Cohen, Jane Cleland-Huang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Journal ref ICSE-SEIS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.00402 2024-02-02 cs.CL cs.AI 62%

Investigating Bias Representations in Llama 2 Chat via Activation Steering

Dawn Lu, Nina Rimsky

专题命中 AI治理与伦理 :RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.16088 2024-01-30 cs.LG cs.CY 62%

Fairness in Algorithmic Recourse Through the Lens of Substantive Equality of Opportunity

Andrew Bell, Joao Fonseca, Carlo Abrate, Francesco Bonchi, Julia Stoyanovich

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10310 2024-01-22 cs.LG cs.AI cs.CC 62%

Mathematical Algorithm Design for Deep Learning under Societal and Judicial Constraints: The Algorithmic Transparency Requirement

Holger Boche, Adalbert Fono, Gitta Kutyniok

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.09473 2024-01-19 cs.CY cs.AI 62%

Business and ethical concerns in domestic Conversational Generative AI-empowered multi-robot systems

Rebekah Rousi, Hooman Samani, Niko Mäkitalo, Ville Vakkuri, Simo Linkola, Kai-Kristian Kemell, Paulius Daubaris, Ilenia Fronza, Tommi Mikkonen, Pekka Abrahamsson

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments 15 pages, 4 figures, International Conference on Software Business

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.06709 2024-01-15 cs.CL cs.AI 62%

Reliability Analysis of Psychological Concept Extraction and Classification in User-penned Text

Muskan Garg, MSVPJ Sathvik, Amrit Chadha, Shaina Raza, Sunghwan Sohn

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏