arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1852 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1852 篇

2505.02174 2025-05-06 cs.CY 57%

AI Governance in the GCC States: A Comparative Analysis of National AI Strategies

Mohammad Rashed Albous, Odeh Rashed Al-Jayyousi, Melodena Stephens

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments 33 pages,6 figures, 11 tables

Journal ref Journal of Artificial Intelligence Research, 82, 2389-2422 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00841 2025-05-05 cs.CR cs.AI 57%

From Texts to Shields: Convergence of Large Language Models and Cybersecurity

Tao Li, Ya-Ting Yang, Yunian Pan, Quanyan Zhu

机构 * New York University(纽约大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21400 2025-05-01 econ.GN cs.CL q-fin.EC 57%

Who Gets the Callback? Generative AI and Gender Bias

Sugat Chaturvedi, Rochana Chaturvedi

机构 * Ahmedabad University(阿赫迈达巴大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20215 2025-04-30 cs.CY cs.HC econ.GN q-fin.EC 57%

Exploring AI-powered Digital Innovations from A Transnational Governance Perspective: Implications for Market Acceptance and Digital Accountability Accountability

Claire Li, David Peter Wallis Freeborn

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Journal ref Proceedings of the UK Academy for Information Systems Conference 2025, Newcastle, UK. UKAIS

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08260 2025-04-18 cs.CL 57%

Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare

Yonchanok Khaokaew, Flora D. Salim, Andreas Züfle, Hao Xue, Taylor Anderson, C. Raina MacIntyre, Matthew Scotch, David J Heslop

机构 * University of New South Wales(新南威尔士大学) Emory University(埃默里大学) George Mason University(乔治梅森大学) Arizona State University(亚利桑那州立大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19363 2025-04-11 cs.CL 57%

Expressivity and Speech Synthesis

Andreas Triantafyllopoulos, Björn W. Schuller

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments Published in Oxford Handbook of Expressivity in Language (in press)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05610 2025-04-09 cs.LG 57%

Fairness in Machine Learning-based Hand Load Estimation: A Case Study on Load Carriage Tasks

Arafat Rahman, Sol Lim, Seokhyun Chung

机构 * University of Virginia(弗吉尼亚大学) Virginia Polytechnic Institute and State University(弗吉尼亚理工大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02874 2025-04-07 cs.CL 57%

TheBlueScrubs-v1, a comprehensive curated medical dataset derived from the internet

Luis Felipe, Carlos Garcia, Issam El Naqa, Monique Shotande, Aakash Tripathi, Vivek Rudrapatna, Ghulam Rasool, Danielle Bitterman, Gilmer Valdes

机构 * Moffitt Cancer Center(莫菲特癌症中心) UCSF(加州大学旧金山分校) The Blue Scrubs Harvard Medical School(哈佛医学院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

Comments 22 pages, 8 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23056 2025-04-07 cs.CY 57%

Achieving Socio-Economic Parity through the Lens of EU AI Act

Arjun Roy, Stavroula Rizou, Symeon Papadopoulos, Eirini Ntoutsi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02074 2025-04-04 cs.HC cs.AI 57%

Trapped by Expectations: Functional Fixedness in LLM-Enabled Chat Search

Jiqun Liu, Jamshed Karimnazarov, Ryen W. White

机构 * The University of Oklahoma(俄克拉荷马大学) Microsoft Research(微软研究院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23566 2025-04-01 cs.CL 57%

When LLM Therapists Become Salespeople: Evaluating Large Language Models for Ethical Motivational Interviewing

Haein Kong, Seonghyeon Moon

机构 * Rutgers University(罗格斯大学) Roblox(罗布乐思公司)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20414 2025-03-27 cs.LG 57%

Active Data Sampling and Generation for Bias Remediation

Antonio Maratea, Rita Perna

机构 * University of Naples Parthenope(那不勒斯 Parthenope 大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09630 2025-03-25 cs.CY cs.HC 57%

What does AI consider praiseworthy?

Andrew J. Peterson

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments Forthcoming in AI and Ethics

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16436 2025-03-24 cs.HC cs.AI cs.RO 57%

Enhancing Human-Robot Collaboration through Existing Guidelines: A Case Study Approach

Yutaka Matsubara, Akihisa Morikawa, Daichi Mizuguchi, Kiyoshi Fujiwara

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09488 2025-03-21 eess.SY cs.LG cs.SY 57%

Intelligent Agricultural Greenhouse Control System Based on Internet of Things and Machine Learning

Cangqing Wang, Jiangchuan Gong

机构 * Boston University(波士顿大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12367 2025-03-18 cs.LG physics.ao-ph 57%

Integrating mobile and fixed monitoring data for high-resolution PM2.5 mapping using machine learning

Rui Xu, Dawen Yao, Yuzhuang Pian, Ruhui Cao, Yixin Fu, Xinru Yang, Ting Gan, Yonghong Liu

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07496 2025-03-14 cs.CY 57%

Securing External Deeper-than-black-box GPAI Evaluations

Alejandro Tlaie, Jimmy Farrell

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09858 2025-03-14 cs.AI cs.GT cs.MA nlin.CD 57%

Media and responsible AI governance: a game-theoretic and LLM analysis

Nataliya Balabanova, Adeela Bashir, Paolo Bova, Alessio Buscemi, Theodor Cimpeanu, Henrique Correia da Fonseca, Alessandro Di Stefano, Manh Hong Duong, Elias Fernandez Domingos, Antonio Fernandes, The Anh Han, Marcus Krellner, Ndidi Bianca Ogbo, Simon T. Powers, Daniele Proverbio, Fernando P. Santos, Zia Ush Shamszaman, Zhao Song

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08931 2025-03-13 cs.CY 57%

ARCHED: A Human-Centered Framework for Transparent, Responsible, and Collaborative AI-Assisted Instructional Design

Hongming Li, Yizirui Fang, Shan Zhang, Seiyon M. Lee, Yiming Wang, Mark Trexler, Anthony F. Botelho

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

Comments Accepted to the iRAISE Workshop at AAAI 2025. To be published in PMLR Volume 273

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08849 2025-03-11 cs.AI 57%

Path To Gain Functional Transparency In Artificial Intelligence With Meaningful Explainability

Md. Tanzib Hosain, Mehedi Hasan Anik, Sadman Rafi, Rana Tabassum, Khaleque Insia, Md. Mehrab Siddiky

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments Hosain, M. T., Anik, M. H., Rafi, S., Tabassum, R., Insia, K., & Sıddıky, M. M. (2023). Path to gain functional transparency in artificial intelligence with meaningful explainability. Journal of Metaverse, 3(2), 166-180

Journal ref Journal of Metaverse, 3(2), 166-180 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02865 2025-03-06 cs.CL 57%

FairSense-AI: Responsible AI Meets Sustainability

Shaina Raza, Mukund Sayeeganesh Chettiar, Matin Yousefabadi, Tahniat Khan, Marcelo Lotif

机构 * Vector Institute for Artificial Intelligence(向量人工智能研究院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05364 2025-03-06 cs.CR cs.AI 57%

Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models

Kalyan Nakka, Jimmy Dani, Nitesh Saxena

机构 * Texas A&M University(德克萨斯农工大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments 26 pages, 31 figures and 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.21250 2025-03-03 cs.AI 57%

Towards Developing Ethical Reasoners: Integrating Probabilistic Reasoning and Decision-Making for Complex AI Systems

Nijesh Upreti, Jessica Ciupa, Vaishak Belle

机构 * The University of Edinburgh(爱丁堡大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18359 2025-02-26 cs.CY 57%

Responsible AI Agents

Deven R. Desai, Mark O. Riedl

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11031 2025-02-18 cs.LG 57%

A Critical Review of Predominant Bias in Neural Networks

Jiazhi Li, Mahyar Khayatkhoei, Jiageng Zhu, Hanchen Xie, Mohamed E. Hussein, Wael AbdAlmageed

机构 * Clemson University(克莱姆森大学) USC Information Sciences Institute(南加州大学信息科学研究所) USC Ming Hsieh Department of Electrical and Computer Engineering(南加州大学明·谢电气与计算机工程学院) USC Thomas Lord Department of Computer Science(南加州大学托马斯·罗德计算机科学学院) Alexandria University(亚历山大大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

Comments 31 pages, 8 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09628 2025-02-18 cs.AI 57%

Artificial Intelligence-Driven Clinical Decision Support Systems

Muhammet Alkan, Idris Zakariyya, Samuel Leighton, Kaushik Bhargav Sivangi, Christos Anagnostopoulos, Fani Deligianni

机构 * University of Glasgow(格拉斯哥大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

Comments Added acknowledgements for the corresponding author, updated Figure 4, 5 & 6

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10400 2025-02-18 cs.CL 57%

Self-Reflection Makes Large Language Models Safer, Less Biased, and Ideologically Neutral

Fengyuan Liu, Nouar AlDahoul, Gregory Eady, Yasir Zaki, Talal Rahwan

机构 * New York University Abu Dhabi(纽约大学阿布扎比分校) University of Copenhagen(哥本哈根大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06472 2025-02-14 cs.RO cs.AI cs.HC 57%

Enabling Novel Mission Operations and Interactions with ROSA: The Robot Operating System Agent

Rob Royce, Marcel Kaufmann, Jonathan Becktor, Sangwoo Moon, Kalind Carpenter, Kai Pak, Amanda Towler, Rohan Thakker, Shehryar Khattak

机构 * NASA Jet Propulsion Laboratory(美国国家航空航天局喷气推进实验室) California Institute of Technology(加州理工学院)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

Comments Preprint. Accepted at IEEE Aerospace Conference 2025, 16 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14040 2025-02-10 cs.CY 57%

Global Perspectives of AI Risks and Harms: Analyzing the Negative Impacts of AI Technologies as Prioritized by News Media

Mowafak Allaham, Kimon Kieslich, Nicholas Diakopoulos

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03472 2025-02-07 cs.CY 57%

Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences

Wesley Pasfield

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

Comments Presented at the 2nd Workshop on Regulatable ML at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏