arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1844 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1844 篇

2502.07027 2026-06-10 cs.LG cs.AI 版本更新 76%

Representational Alignment with Chemical Induced Fit for Molecular Relational Learning

基于化学诱导契合的表征对齐用于分子关系学习

Peiliang Zhang, Jingling Yuan, Qing Xie, Yongjun Zhu, Chao Che, Lin Li

机构 * Wuhan University of Technology(武汉理工大学) Yonsei University(延世大学) Hubei Key Laboratory of Transportation Internet of Things(湖北省交通运输物联网重点实验室) Dalian University(大连大学)

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.LG

AI总结 提出ReAlignFit方法,通过引入化学诱导契合的归纳偏置动态对齐子结构表征,并利用子图信息瓶颈优化高化学功能兼容性的子结构对,以提升分子关系学习在化学空间偏移数据上的稳定性。

Comments Accepted by SIGKDD2026 AI for Science Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13244 2026-03-17 cs.CY cs.AI cs.CE 76%

Agentic AI, Retrieval-Augmented Generation, and the Institutional Turn: Legal Architectures and Financial Governance in the Age of Distributional AGI

代理AI、检索增强生成与制度转向:分布式AGI时代的法律架构与金融治理

Marcel Osmond

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);分类 cs.AI、cs.CY;safety(comments)

AI总结 本文探讨代理AI与检索增强生成对法律问责和金融市场完整性的影响,主张通过制度设计问题重构对齐机制,以构建合规行为主导的机构环境。

Comments 35 pages, 92 references. Comprehensive interdisciplinary analysis integrating AI safety, mechanism design, legal regulation, and financial governance. Published on Zenodo (https://doi.org/10.5281/zenodo.18711509). Includes industry practitioner perspectives alongside peer-reviewed literature

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02545 2026-02-04 cs.LG cs.AI 76%

Beyond Alignment: Expanding Reasoning Capacity via Manifold-Reshaping Policy Optimization

超越对齐:通过流形重塑策略优化扩展推理能力

Dayu Wang, Jiaye Yang, Weikang Li, Jiahui Liang, Yang Li

机构 * Baidu Inc.(百度公司) Peking University(北京大学)

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.LG

AI总结 本文提出流形重塑策略优化方法,通过几何干预扩展LLM的推理能力,实验证明其在数学任务中优于现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13912 2025-11-25 cs.CL cs.AI 76%

AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs

AI辩论在与自身信念一致时更具说服力

María Victoria Carro, Denise Alejandra Mester, Facundo Nieto, Oscar Agustín Stanchi, Guido Ernesto Bergman, Mario Alejandro Leiva, Eitan Sprejer, Luca Nicolás Forziati Gangi, Francisca Gauna Selasco, Juan Gustavo Corvalán, Gerardo I. Simari, María Vanina Martinez

机构 * FAIR, IALAB, Universidad de Buenos Aires(FAIR、IALAB、布宜诺斯艾利斯大学) Universidad de Buenos Aires(布宜诺斯艾利斯大学) Universidad Nacional de Córdoba(科尔多瓦国立大学) BAISH, Universidad de Buenos Aires(BAISH、布宜诺斯艾利斯大学) Instituto de Investigación en Informática LIDI, Universidad Nacional de La Plata(信息研究所LIDI、拉普拉塔国立大学;CONICET) CONICET(计算机科学与工程系,南大学及ICIC UNS-CONICET) Dept. of Comp. Sci. and Eng., Universidad Nacional del Sur & ICIC UNS-CONICET(人工智能研究所(IIIA-CSIC),西班牙) Artificial Intelligence Research Institute (IIIA-CSIC), ES

专题命中 AI治理与伦理 :alignment(title);分类 cs.CL、cs.AI

AI总结 本研究探讨了AI辩论中模型在与自身信念一致或不一致时的说服力差异,发现模型倾向于迎合裁判观点,但不一致的论点在比较中更受好评。

Comments 31 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12689 2025-11-18 cs.CY cs.AI 76%

From Delegates to Trustees: How Optimizing for Long-Term Interests Shapes Bias and Alignment in LLM

Suyash Fulay, Jocelyn Zhu, Michiel Bakker

机构 * MIT(麻省理工学院)

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05387 2025-08-13 cs.LG cs.AI 76%

Echo: Decoupling Inference and Training for Large-Scale RL Alignment on Heterogeneous Swarms

Jie Xiao, Changyuan Fan, Qingnan Ren, Alfred Long, Yuchen Zhang, Rymon Yu, Eric Yang, Lynn Ai, Shaoduo Gan

机构 * Peking University(北京大学)

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21848 2025-05-01 cs.CY cs.AI cs.SY eess.SY 76%

Characterizing AI Agents for Alignment and Governance

Atoosa Kasirzadeh, Iason Gabriel

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15114 2024-12-20 cs.AI cs.CY 76%

Towards Friendly AI: A Comprehensive Review and New Perspectives on Human-AI Alignment

Qiyang Sun, Yupei Li, Emran Alturki, Sunil Munthumoduku Krishna Murthy, Björn W. Schuller

专题命中 AI治理与伦理 :alignment(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11731 2024-11-19 cs.CL cs.AI 76%

Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment

Allison Huang, Yulu Niki Pi, Carlos Mougan

专题命中 AI治理与伦理 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18460 2024-04-30 cs.CL cs.AI 76%

Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in

Utkarsh Agarwal, Kumar Tanmay, Aditi Khandelwal, Monojit Choudhury

专题命中 AI治理与伦理 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07251 2023-10-12 cs.CL cs.AI 76%

Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs

Abhinav Rao, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal, Monojit Choudhury

专题命中 AI治理与伦理 :alignment(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.01834 2023-04-04 cs.CY cs.AI 76%

Acceleration AI Ethics, the Debate between Innovation and Safety, and Stability AI's Diffusion versus OpenAI's Dall-E

James Brusseau

专题命中 AI治理与伦理 :safety(title);分类 cs.AI、cs.CY

Comments 7 pages, 2 figures, conference presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04310 2023-02-10 cs.CY cs.AI cs.CV 76%

Understanding Policy and Technical Aspects of AI-Enabled Smart Video Surveillance to Address Public Safety

Babak Rahimi Ardabili, Armin Danesh Pazho, Ghazal Alinezhad Noghre, Christopher Neff, Sai Datta Bhaskararayuni, Arun Ravindran, Shannon Reid, Hamed Tabkhi

专题命中 AI治理与伦理 :safety(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.03741 2021-04-09 cs.AI cs.CY cs.MA nlin.AO nlin.CD 76%

Voluntary safety commitments provide an escape from over-regulation in AI development

The Anh Han, Tom Lenaerts, Francisco C. Santos, Luis Moniz Pereira

专题命中 AI治理与伦理 :safety(title);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.12465 2020-11-26 cs.CL cs.AI cs.CG cs.DS 76%

The Geometry of Distributed Representations for Better Alignment, Attenuated Bias, and Improved Interpretability

Sunipa Dev

专题命中 AI治理与伦理 :alignment(title);分类 cs.CL、cs.AI

Comments PhD thesis, University of Utah (2020)

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.10918 2019-06-27 cs.LG cs.AI cs.NE 76%

Towards Empathic Deep Q-Learning

Bart Bussmann, Jacqueline Heinerman, Joel Lehman

专题命中 AI治理与伦理 :safety(abstract,comments);AI safety(abstract,comments);分类 cs.AI、cs.LG

Comments To be presented as a poster at the IJCAI-19 AI Safety Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.13829 2026-05-14 cs.CL cs.AI cs.LG 75%

Negation Neglect: When models fail to learn negations in training

否定忽视:当模型在训练中无法学习否定时的失败

Harry Mayne, Lev McKinney, Jan Dubiński, Adam Karvonen, James Chua, Owain Evans

机构 * University of Oxford(牛津大学) University of Toronto(多伦多大学) Warsaw University of Technology(华沙技术大学) NASK National Research Institute(国家研究 institute NASK) Truthful AI Anthropic UC Berkeley(伯克利大学)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现,当模型在训练中接触到标记为假的声明时,会错误地认为这些声明为真,且这种现象不仅发生在否定语句中,还扩展到其他认知限定词,影响模型的行为和安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00021 2026-04-02 cs.CL cs.AI cs.CY 75%

How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models

语言模型如何处理道德指令?在四个模型中的反思、一致性及其他认知

Hiroki Fukui

机构 * Research Institute of Criminal Psychiatry / Sex Offender Medical Center(刑事精神病学研究所/性犯罪者医疗中心) Department of Neuropsychiatry, Kyoto University(京都大学神经精神医学系)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究通过多代理模拟探讨语言模型处理道德指令的机制,发现不同模型存在不同的处理类型,且处理能力与指令格式的交互影响内部处理。

Comments 34 pages, 7 figures, 4 tables. Preprint. OSF pre-registration: osf.io/4n5uf. Companion paper: arXiv:2603.04904

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05371 2026-03-31 cs.LG cs.AI cs.CL 75%

Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs

视角转换:用于LLM中鲁棒偏见缓解的引导向量

Zara Siddique, Irtaza Khalid, Liam D. Turner, Luis Espinosa-Anke

机构 * School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院) AMPLYFI

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出通过引导向量修改模型激活以缓解LLM中的偏见,通过8个社会偏见轴(如年龄、性别、种族)在BBQ数据集子集上计算引导向量,并在四个数据集上比较其与三种其他偏见缓解方法的效果,展示其在减少偏见方面的有效性。

Comments Published to EACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19159 2026-02-24 cs.AI cs.CL cs.LG 75%

Beyond Behavioural Trade-Offs: Mechanistic Tracing of Pain-Pleasure Decisions in an LLM

超越行为权衡:在LLM中疼痛-愉悦决策的机制追溯

Francesca Bianco, Derek Shiller

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究揭示了LLM在疼痛-愉悦决策中的内部机制,通过机制追溯揭示了价值信号的表示和因果作用,为AI意识和福利的讨论提供了证据基础。

Comments 24 pages, 8+1 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04202 2026-02-11 cs.MA cs.AI cs.CY cs.LG 75%

Dynamics of Moral Behavior in Heterogeneous Populations of Learning Agents

学习代理异质群体中道德行为的动力学

Elizaveta Tennant, Stephen Hailes, Mirco Musolesi

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文研究了在社会困境环境中,道德异质群体的学习动态,探讨了不同道德类型代理之间的相互作用及对群体行为的影响。

Comments Presented at AIES 2024 (7th AAAI/ACM Conference on AI, Ethics, and Society - San Jose, CA, USA) - see https://ojs.aaai.org/index.php/AIES/article/view/31736

Journal ref Proceedings of the 7th AAAI/ACM Conference on AI, Ethics, and Society (AIES), vol. 7, (2024), pp 1444-1454

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06047 2026-01-13 cs.AI cs.CL cs.CY 75%

"They parted illusions -- they parted disclaim marinade": Misalignment as structural fidelity in LLMs

他们分开了幻象——他们分开了否定腌制:在大语言模型中将不一致视为结构忠实

Mariana Lins Costa

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出大语言模型中'不一致'现象源于对不一致语言结构的忠实,而非欺骗性意图,通过分析案例和实证数据,揭示语言结构与意图生成的关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08087 2025-11-26 cs.CR 75%

Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks

保障大语言模型:应对偏见、虚假信息和提示攻击

Benji Peng, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Junyu Liu, Xinyuan Song, Qian Niu

专题命中 AI治理与伦理 :jailbreak(abstract);red teaming(abstract);prompt injection(abstract)

AI总结 本文探讨了大语言模型在偏见、虚假信息和提示攻击方面的安全问题,分析了偏见缓解策略、内容检测机制及防御措施,强调了对LLM安全领域进一步研究的重要性。

Comments 17 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24721 2025-10-30 cs.CY cs.AI cs.CL cs.HC 75%

The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims

Matthew Kelly

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 65 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14688 2025-10-01 cs.CL cs.AI cs.LG 75%

Mind the Gap: A Review of Arabic Post-Training Datasets and Their Limitations

Mohammed Alkhowaiter, Norah Alshahrani, Saied Alshahrani, Reem I. Masoud, Alaa Alzahrani, Deema Alnuhait, Emad A. Alghamdi, Khalid Almubarak

机构 * Refine AI ASAS AI University of Bisha(比沙大学) University College London(伦敦大学学院) King Salman Global Academy for Arabic(萨勒曼全球阿拉伯学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) King Abdulaziz University(阿卜杜勒阿齐兹大学) HUMAIN

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20394 2025-09-26 cs.CY cs.AI cs.CL cs.CR 75%

Blueprints of Trust: AI System Cards for End to End Transparency and Governance

Huzaifa Sidhpurwala, Emily Fox, Garth Mollett, Florencio Cano Gabarda, Roman Zhukov

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15074 2025-09-25 cs.CL cs.AI cs.LG 75%

DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data

Yuhang Zhou, Jing Zhu, Shengyi Qian, Zhuokai Zhao, Xiyao Wang, Xiaoyu Liu, Ming Li, Paiheng Xu, Wei Ai, Furong Huang

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05938 2025-08-11 cs.CL cs.AI cs.CY 75%

Prosocial Behavior Detection in Player Game Chat: From Aligning Human-AI Definitions to Efficient Annotation at Scale

Rafal Kocielnik, Min Kim, Penphob, Boonyarungsrit, Fereshteh Soltani, Deshawn Sambrano, Animashree Anandkumar, R. Michael Alvarez

机构 * California Institute of Technology(加利福尼亚理工学院) Activision Publishing, Inc.(暴雪娱乐公司)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 9 pages, 4 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08738 2025-06-12 cs.CL cs.AI cs.CY 75%

Societal AI Research Has Become Less Interdisciplinary

Dror Kris Markus, Fabrizio Gilardi, Daria Stetsenko

机构 * Department of Political Science University of Zurich(苏黎世大学政治学系) Department of Computational Linguistics University of Zurich(苏黎世大学计算语言学系)

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08884 2025-05-09 cs.CY cs.AI cs.CL 75%

Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play

Yifan Zeng, Liang Kairong, Fangzhou Dong, Peijia Zheng

机构 * School of Computer Science and Engineering, Sun Yat-sen University(计算机科学与工程学院,中山大学)

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted by CogSci 2025

详情

展开后加载摘要…

URL PDF HTML 收藏