arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2505.13973 2025-05-21 cs.CL cs.AI cs.CV 62%

Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models

Wenhui Zhu, Xuanzhao Dong, Xin Li, Peijie Qiu, Xiwen Chen, Abolfazl Razi, Aris Sotiras, Yi Su, Yalin Wang

机构 * Arizona State University(亚利桑那州立大学) Clemson University(克莱姆森大学) Washington University in St.Louis(华盛顿大学圣路易斯分校) Banner Alzheimer’s Institute(Banner阿尔茨海默病研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12625 2025-05-20 cs.CL cs.CR cs.LG 62%

R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model

Ali Naseh, Harsh Chaudhari, Jaechul Roh, Mingshi Wu, Alina Oprea, Amir Houmansadr

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Northeastern University(东北大学) GFW Report(GFW报告)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15508 2025-05-16 cs.CL cs.AI 62%

Compensate Quantization Errors+: Quantized Models Are Inquisitive Learners

Yifei Gao, Jie Ou, Lei Wang, Jun Cheng, Mengchu Zhou

机构 * Yifei Gao ∗ , Jie Ou ∗ , Lei Wang † † {\dagger} † , Jun Cheng, and Mengchu Zhou ∗(作者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments Effecient Quantization Methods for LLMs

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09576 2025-05-15 cs.CY cs.AI 62%

Ethics and Persuasion in Reinforcement Learning from Human Feedback: A Procedural Rhetorical Approach

Shannon Lodoen, Alexi Orchard

专题命中 AI治理与伦理 :RLHF(abstract);分类 cs.AI、cs.CY

Comments 10 pages, 1 figure, Accepted version

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08106 2025-05-14 cs.CL cs.AI 62%

Are LLMs complicated ethical dilemma analyzers?

Jiashen, Du, Jesse Yao, Allen Liu, Zhekai Zhang

机构 * Department of Computer Science, University of California, Berkeley(加州大学伯克利分校计算机科学系)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments CS194-280 Advanced LLM Agents project. Project page: https://github.com/ALT-JS/ethicaLLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08064 2025-05-14 cs.HC cs.AI cs.CY 62%

Justified Evidence Collection for Argument-based AI Fairness Assurance

Alpay Sabuncuoglu, Christopher Burr, Carsten Maple

机构 * The Alan Turing Institute United Kingdom(阿尔法·图灵研究所(英国)) University of Warwick United Kingdom(沃里克大学(英国))

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments The paper is accepted for ACM Conference on Fairness, Accountability, and Transparency (ACM FAccT '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07875 2025-05-14 cs.CY cs.AI 62%

Getting Ready for the EU AI Act in Healthcare. A call for Sustainable AI Development and Deployment

John Brandt Brodersen, Ilaria Amelia Caggiano, Pedro Kringen, Vince Istvan Madai, Walter Osika, Giovanni Sartor, Ellen Svensson, Magnus Westerlund, Roberto V. Zicari

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 8 pages, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06326 2025-05-13 cs.CY cs.AI 62%

Enterprise Architecture as a Dynamic Capability for Scalable and Sustainable Generative AI adoption: Bridging Innovation and Governance in Large Organisations

Alexander Ettinger

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 82 pages excluding appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.14362 2025-05-12 cs.HC cs.AI cs.CY 62%

The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health Support

Inhwa Song, Sachin R. Pendse, Neha Kumar, Munmun De Choudhury

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments The first two authors contributed equally to this work; typos corrected and post-review revisions incorporated

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05197 2025-05-09 cs.AI cs.CY 62%

Societal and technological progress as sewing an ever-growing, ever-changing, patchy, and polychrome quilt

Joel Z. Leibo, Alexander Sasha Vezhnevets, William A. Cunningham, Sébastien Krier, Manfred Diaz, Simon Osindero

机构 * Google DeepMind(谷歌DeepMind) University of Toronto(多伦多大学) Mila - Québec AI Institute(魁北克AI研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15941 2025-05-06 cs.CL cs.AI 62%

FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity

Fanny Jourdan, Yannick Chevalier, Cécile Favre

机构 * IRT Saint Exupery(IRT圣埃克苏佩里) Université Lumière Lyon 2(里莫大学 Lyon 2) Université Claude Bernard Lyon 1(克劳德·贝尔纳大学 Lyon 1) ERIC(埃里克)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments FAccT 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04785 2025-05-06 cs.CL cs.CY 62%

Mapping Trustworthiness in Large Language Models: A Bibliometric Analysis Bridging Theory to Practice

José Siqueira de Cerqueira, Kai-Kristian Kemell, Rebekah Rousi, Nannan Xi, Juho Hamari, Pekka Abrahamsson

机构 * Tampere University(塔尔基耶大学) University of Vaasa(瓦萨大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00339 2025-05-02 cs.CL cs.AI 62%

Enhancing AI-Driven Education: Integrating Cognitive Frameworks, Linguistic Feedback Analysis, and Ethical Considerations for Improved Content Generation

Antoun Yaacoub, Sansiri Tarnpradab, Phattara Khumprom, Zainab Assaghir, Lionel Prevost, Jérôme Da-Rugna

机构 * Learning, Data and Robotics (LDR) ESIEA Lab(学习、数据与机器人(LDR)ESIEA实验室) Department of Computer Engineering(计算机工程系) Graduate School of Management and Innovation(管理与创新研究生院) Faculty of Science(科学学院) Lebanese University(黎巴嫩大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

Comments This article will be presented in IJCNN 2025 "AI Innovations for Education: Transforming Teaching and Learning through Cutting-Edge Technologies" workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09495 2025-05-01 cs.LG cs.AI 62%

FADE: Towards Fairness-aware Generation for Domain Generalization via Classifier-Guided Score-based Diffusion Models

Yujie Lin, Dong Li, Minglai Shao, Guihong Wan, Chen Zhao

机构 * School of New Media and Communication, Tianjin University(天津大学新媒体与传播学院) School of Informatics, Xiamen University(厦门大学信息学院) Department of Computer Science, Baylor University(贝勒大学计算机科学系) Departments of Biostatistics and Epidemiology, Harvard University(哈佛大学生物统计学与流行病学系)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19255 2025-04-29 cs.AI cs.CY 62%

The Convergent Ethics of AI? Analyzing Moral Foundation Priorities in Large Language Models with a Multi-Framework Approach

Chad Coleman, W. Russell Neuman, Ali Dasdan, Safinah Ali, Manan Shah

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 25 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17401 2025-04-29 cs.CY cs.AI cs.HC 62%

AIJIM: A Scalable Model for Real-Time AI in Environmental Journalism

Torsten Tiltack

机构 * Torsten Tiltack(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 22 pages, 10 figures, 5 tables. Keywords: Artificial Intelligence, Environmental Journalism, Real-Time Reporting, Vision Transformers, Image Recognition, Crowdsourced Validation, GPT-4, Automated News Generation, GIS Integration, Data Privacy Compliance, Explainable AI (XAI), AI Ethics, Sustainable Development

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04476 2025-04-28 cs.CY cs.AI 62%

The Moral Mind(s) of Large Language Models

Avner Seror

机构 * Aix Marseille Univ, CNRS, AMSE, Marseille, France(阿维尼翁-马赛大学,国家科学研究中心,AMSE,马赛,法国)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16948 2025-04-25 cs.CY cs.AI cs.ET 62%

Intrinsic Barriers to Explaining Deep Foundation Models

Zhen Tan, Huan Liu

机构 * Arizona State University(亚利桑那州立大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16622 2025-04-24 cs.AI cs.CY 62%

Cognitive Silicon: An Architectural Blueprint for Post-Industrial Computing Systems

Christoforus Yoga Haryanto, Emily Lomempow

机构 * ZipThought

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Working Paper, 37 pages, 1 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13957 2025-04-22 cs.CY cs.AI cs.CR 62%

Naming is framing: How cybersecurity's language problems are repeating in AI governance

Lianne Potter

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments 20 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09946 2025-04-21 cs.CY cs.CL 62%

Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Qian Wang, Zhanzhi Lou, Zhenheng Tang, Nuo Chen, Xuandong Zhao, Wenxuan Zhang, Dawn Song, Bingsheng He

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12358 2025-04-18 cs.CY cs.AI physics.soc-ph 62%

Towards an AI Observatory for the Nuclear Sector: A tool for anticipatory governance

Aditi Verma, Elizabeth Williams

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments Presented at the Sociotechnical AI Governance Workshop at CHI 2025, Yokohama

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11501 2025-04-17 cs.CY cs.AI 62%

A Framework for the Private Governance of Frontier Artificial Intelligence

Dean W. Ball

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00041 2025-04-15 eess.SP cs.AI cs.LG 62%

Needles in Needle Stacks: Meaningful Clinical Information Buried in Noisy Waveform Data

Sujay Nagaraj, Andrew J. Goodwin, Dmytro Lopushanskyy, Danny Eytan, Robert W. Greer, Sebastian D. Goodfellow, Azadeh Assadi, Anand Jayarajan, Anna Goldenberg, Mjaye L. Mazwi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

Comments Machine Learning For Health Care 2024 (MLHC)

Journal ref PMLR (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08640 2025-04-14 cs.AI cs.CY cs.GT nlin.CD 62%

Do LLMs trust AI regulation? Emerging behaviour of game-theoretic LLM agents

Alessio Buscemi, Daniele Proverbio, Paolo Bova, Nataliya Balabanova, Adeela Bashir, Theodor Cimpeanu, Henrique Correia da Fonseca, Manh Hong Duong, Elias Fernandez Domingos, Antonio M. Fernandes, Marcus Krellner, Ndidi Bianca Ogbo, Simon T. Powers, Fernando P. Santos, Zia Ush Shamszaman, Zhao Song, Alessandro Di Stefano, The Anh Han

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07516 2025-04-11 cs.CY cs.AI cs.HC 62%

Enhancements for Developing a Comprehensive AI Fairness Assessment Standard

Avinash Agarwal, Mayashankar Kumar, Manisha J. Nene

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

Comments 5 pages. Published in 2025 17th International Conference on COMmunication Systems and NETworks (COMSNETS). Access: https://ieeexplore.ieee.org/abstract/document/10885551

Journal ref 2025 17th International Conference on COMmunication Systems and NETworks (COMSNETS), Bengaluru, India, 2025, pp. 1216-1220

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07118 2025-04-11 cs.CY cs.AI cs.ET 62%

Sacred or Secular? Religious Bias in AI-Generated Financial Advice

Muhammad Salar Khan, Hamza Umer

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12848 2025-04-10 cs.CY cs.AI cs.SI 62%

ClarityEthic: Explainable Moral Judgment Utilizing Contrastive Ethical Insights from Large Language Models

Yuxi Sun, Wei Gao, Jing Ma, Hongzhan Lin, Ziyang Luo, Wenxuan Zhang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

Comments We have noticed that this version of our experiment and method description isn't quite complete or accurate. To make sure we present our best work, we think it would be a good idea to withdraw the manuscript for now and take some time to revise and reformat it

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02917 2025-04-07 cs.CL cs.AI 62%

Bias in Large Language Models Across Clinical Applications: A Systematic Review

Thanathip Suenghataiphorn, Narisara Tribuddharat, Pojsakorn Danpanichkul, Narathorn Kulthamrongsri

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01810 2025-04-03 cs.CY cs.AI 62%

Propaganda is all you need

Paul Kronlund-Drouault

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏