arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1844 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1844 篇

2603.18788 2026-03-25 cs.CL cs.AI 73%

Mi:dm K 2.5 Pro

KT Tech innovation Group

专题命中 AI治理与伦理 :safety(abstract);harmlessness(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19144 2026-03-20 cs.CL cs.AI 73%

UGID: Unified Graph Isomorphism for Debiasing Large Language Models

UGID: 用于去偏大型语言模型的统一图同构

Zikang Ding, Junchi Yao, Junhao Li, Yi Zhang, Wenbo Jiang, Hongbo Liu, Lijie Hu

机构 * University of Electronic Science and Technology of China(电子科技大学) Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学) South China University of Technology(华南理工大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

AI总结 UGID通过构建Transformer的结构化计算图,在内部表示层面减少模型偏见,通过约束注意力路由和隐藏表示,提升模型行为一致性并保持安全性和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14495 2026-03-17 cs.CY cs.AI 73%

Bridging the Gap in the Responsible AI Divides

弥合负责任AI分歧的鸿沟

Bálint Gyevnár, Atoosa Kasirzadeh

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

AI总结 本文通过分析3550篇论文,探讨AI安全与AI伦理之间的分歧与重叠,提出通过解决关键问题促进负责任的AI治理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13900 2026-03-06 cs.CL cs.AI 73%

Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences

细粒度微调在激活差异中留下明显可读的痕迹

Julian Minder, Clément Dumas, Stewart Slocum, Helena Casademunt, Cameron Holmes, Robert West, Neel Nanda

机构 * EPFL(苏黎世联邦理工学院) Ecole Normale Supérieure Paris-Saclay(巴黎-萨克雷高等师范学校) Université Paris-Saclay(巴黎-萨克雷大学) Harvard University(哈佛大学) MATS

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

AI总结 研究发现狭窄微调会在激活中留下明显痕迹,揭示了微调领域偏见,并警告了使用此类模型进行广泛微调研究的局限性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04742 2026-02-05 cs.CY cs.CL 73%

Inference-Time Reasoning Selectively Reduces Implicit Social Bias in Large Language Models

推理时的推理选择性减少大语言模型中的隐性社会偏见

Molly Apsel, Michael N. Jones

机构 * Cognitive Science Program Department of Psychological & Brain Sciences Indiana University Bloomington(心理与脑科学系认知科学计划印第安纳大学布卢明顿)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.CY

AI总结 该研究通过推理增强减少大语言模型中的隐性社会偏见,揭示推理对公平性评估的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14017 2025-11-19 cs.LG cs.AI 73%

From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs

Erum Mushtaq, Anil Ramakrishna, Satyapriya Krishna, Sattvik Sahai, Prasoon Goyal, Kai-Wei Chang, Tao Zhang, Rahul Gupta

机构 * University of Southern California(南加州大学) Amazon AGI(亚马逊人工智能实验室)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05901 2025-11-14 cs.CL cs.AI 73%

Retrieval-Augmented Generation in Medicine: A Scoping Review of Technical Implementations, Clinical Applications, and Ethical Considerations

Rui Yang, Matthew Yu Heng Wong, Huitao Li, Xin Li, Wentao Zhu, Jingchi Liao, Kunyu Yu, Jonathan Chong Kai Liew, Weihao Xuan, Yingjian Chen, Yuhe Ke, Jasmine Chiat Ling Ong, Douglas Teodoro, Chuan Hong, Daniel Shi Wei Ting, Nan Liu

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23994 2025-11-10 cs.CL cs.AI 73%

Policy-as-Prompt: Turning AI Governance Rules into Guardrails for AI Agents

Gauri Kholkar, Ratinder Ahuja

机构 * Pure Storage

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments Accepted at 3rd Regulatable ML Workshop at NEURIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02895 2025-11-07 cs.CY cs.AI cs.HC physics.soc-ph 73%

A Criminology of Machines

Gian Maria Campedelli

机构 * Fondazione Bruno Kessler(布鲁诺·科斯勒基金会)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments This pre-print is also available at CrimRxiv with DOI: https://doi.org/10.21428/cb6ab371.e3354ce1

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25445 2025-10-30 cs.AI cs.LG 73%

Agentic AI: A Comprehensive Survey of Architectures, Applications, and Future Directions

Mohamad Abou Ali, Fadi Dornaika

机构 * University of the Basque Country(巴斯克大学) Lebanese International University (LIU)(黎巴嫩国际大学) The International University of Beirut(贝鲁特国际大学) IKERBASQUE

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19008 2025-10-23 cs.HC cs.AI cs.LG cs.MA 73%

Plural Voices, Single Agent: Towards Inclusive AI in Multi-User Domestic Spaces

Joydeep Chandra, Satyam Kumar Navneet

机构 * BNRIST, Tsinghua University(北京理工大学、清华大学) Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16555 2025-10-21 cs.AI cs.LG 73%

Urban-R1: Reinforced MLLMs Mitigate Geospatial Biases for Urban General Intelligence

Qiongyan Wang, Xingchen Zou, Yutian Jiang, Haomin Wen, Jiaheng Wei, Qingsong Wen, Yuxuan Liang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Carnegie Mellon University(卡内基梅隆大学) Squirrel Ai Learning

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07887 2025-10-17 cs.CL cs.AI 73%

Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge

Riccardo Cantini, Alessio Orsino, Massimo Ruggiero, Domenico Talia

机构 * University of Calabria(卡利博利亚大学)

专题命中 AI治理与伦理 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Journal ref Cantini, R., Orsino, A., Ruggiero, M., Talia, D. Benchmarking adversarial robustness to bias elicitation in large language models: scalable automated assessment with LLM-as-a-judge. Mach Learn 114, 249 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02444 2025-10-08 cs.CL cs.AI 73%

Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language Models

Haoran Ye, Tianze Zhang, Yuhang Xie, Liyuan Zhang, Yuanyi Ren, Xin Zhang, Guojie Song

机构 * State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University(人工智能通用基础理论国家重点实验室,智能科学与技术学院,北京大学) Yuanpei College, Peking University(元培学院,北京大学) School of Psychological and Cognitive Sciences, Peking University(心理学与认知科学学院,北京大学) Key Laboratory of Machine Perception (Ministry of Education), Peking University(机器感知重点实验室(教育部),北京大学) PKU-Wuhan Institute for Artificial Intelligence(北京大学-武汉人工智能研究院)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04931 2025-09-04 cs.AI cs.CL cs.HC 73%

A Survey on Human-AI Collaboration with Large Foundation Models

Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yanwen Xu, Lei Ding, Jay Mehta, Nathan Grinnell, Li Liu, Sijia Zhong, Devanathan Nallur Gandamani, Xinyi Tang, Rohan Ghosalkar, Celeste Shen, Rachel Shen, Nafisa Hussain, Kesav Ravichandran, James Davis

机构 * University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments Topic and scope refinement

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00085 2025-09-03 cs.CR cs.AI cs.CY 73%

Private, Verifiable, and Auditable AI Systems

Tobin South

机构 * MIT Media Lab(麻省理工学院媒体实验室)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments PhD thesis

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06849 2025-08-12 cs.CY cs.AI cs.HC 73%

Towards Experience-Centered AI: A Framework for Integrating Lived Experience in Design and Development

Sanjana Gautam, Mohit Chandra, Ankolika De, Tatiana Chakravorti, Girik Malik, Munmun De Choudhury

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11502 2025-07-16 cs.CL cs.CE cs.LG 73%

HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong

Sirui Han, Junqi Zhu, Ruiyuan Zhang, Yike Guo

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03050 2025-07-08 cs.CY cs.AI 73%

From Turing to Tomorrow: The UK's Approach to AI Regulation

Oliver Ritchie, Markus Anderljung, Tom Rachman

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments This is a chapter intended for publication in a forthcoming edited volume. It is the version of the author's manuscript prior to acceptance for publication and has not undergone editorial and/or peer review on behalf of the Publisher

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05050 2025-06-04 cs.CL cs.AI 73%

Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models

Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang, Shaohui Mei, Lap-Pui Chau

机构 * Department of Electrical and Electronic Engineering(电子与电气工程系) The Hong Kong Polytechnic University(香港理工大学) School of Electronics and Information(电子与信息学院) Northwestern Polytechnical University(西北工业大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01257 2025-06-03 cs.CL cs.AI 73%

DeepSeek in Healthcare: A Survey of Capabilities, Risks, and Clinical Applications of Open-Source Large Language Models

Jiancheng Ye, Sophie Bronstein, Jiarui Hai, Malak Abu Hashish

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17131 2025-05-26 cs.CL cs.AI stat.ML 73%

Relative Bias: A Comparative Framework for Quantifying Bias in LLMs

Alireza Arbabi, Florian Kerschbaum

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.04756 2025-05-23 cs.CL cs.LG 73%

Red-Teaming for Inducing Societal Bias in Large Language Models

Chu Fei Luo, Ahmad Ghawanmeh, Bharat Bhimshetty, Kashyap Murali, Murli Jadhav, Xiaodan Zhu, Faiza Khan Khattak

机构 * Queen’s University Vector Institute(女王大学向量研究所) SigmaRed Tech.(SigmaRed科技公司) Ernst & Young(埃森哲) Monark Health(Monark健康)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13972 2025-04-22 cs.CY cs.AI 73%

Governance Challenges in Reinforcement Learning from Human Feedback: Evaluator Rationality and Reinforcement Stability

Dana Alsagheer, Abdulrahman Kamal, Mohammad Kamal, Weidong Shi

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09097 2024-12-18 cs.CL cs.AI 73%

Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations

Tarun Raheja, Nilay Pochhi, F. D. C. M. Curie

专题命中 AI治理与伦理 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Comments 16 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14194 2024-10-21 cs.CL cs.AI 73%

Speciesism in Natural Language Processing Research

Masashi Takeshita, Rafal Rzepka

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments This article is a preprint and has not been peer-reviewed. The postprint has been accepted for publication in AI and Ethics. Please cite the final version of the article once it is published

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08323 2024-09-18 cs.CY cs.AI 73%

Mapping the Ethics of Generative AI: A Comprehensive Scoping Review

Thilo Hagendorff

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.04023 2024-08-09 cs.CL cs.AI 73%

Improving Large Language Model (LLM) fidelity through context-aware grounding: A systematic approach to reliability and veracity

Wrick Talukdar, Anjanava Biswas

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments 14 pages

Journal ref World Journal of Advanced Engineering Technology and Sciences, 2023, 10(2), 283-296

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03820 2024-06-14 cs.CY cs.AI cs.HC 73%

False Sense of Security in Explainable Artificial Intelligence (XAI)

Neo Christopher Chung, Hongkyou Chung, Hearim Lee, Lennart Brocki, Hongbeom Chung, George Dyer

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

Comments AI Governance Workshop at the 2024 International Joint Conference on Artificial Intelligence (IJCAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12744 2024-05-13 cs.CL cs.AI 73%

Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches

Pablo Biedma, Xiaoyuan Yi, Linus Huang, Maosong Sun, Xing Xie

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments 16 pages, work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏