arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1844 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1844 篇

2405.01886 2024-05-06 cs.CL cs.AI 73%

Aloe: A Family of Fine-tuned Open Healthcare LLMs

Ashwin Kumar Gururajan, Enrique Lopez-Cuena, Jordi Bayarri-Planas, Adrian Tormos, Daniel Hinjos, Pablo Bernabeu-Perez, Anna Arias-Duart, Pablo Agustin Martin-Torres, Lucia Urcelay-Ganzabal, Marta Gonzalez-Mallo, Sergio Alvarez-Napagao, Eduard Ayguadé-Parra, Ulises Cortés Dario Garcia-Gasulla

专题命中 AI治理与伦理 :alignment(abstract);red teaming(abstract);分类 cs.CL、cs.AI

Comments Five appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.19076 2024-05-01 cs.CY cs.AI 73%

Who Followed the Blueprint? Analyzing the Responses of U.S. Federal Agencies to the Blueprint for an AI Bill of Rights

Darren Lage, Riley Pruitt, Jason Ross Arnold

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.06263 2023-11-14 cs.CY cs.AI 73%

No Trust without regulation!

François Terrier

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.15370 2023-06-07 cs.CY cs.AI 73%

A Principles-based Ethics Assurance Argument Pattern for AI and Autonomous Systems

Zoe Porter, Ibrahim Habli, John McDermid, Marten Kaas

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.CY

Journal ref AI and Ethics 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.11163 2023-04-25 cs.CY cs.CL 73%

ChatGPT, Large Language Technologies, and the Bumpy Road of Benefiting Humanity

Atoosa Kasirzadeh

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.CY

Comments As part of a series on Dailynous : "Philosophers on next-generation large language models"

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.04776 2023-01-04 cs.LG cs.AI cs.CV cs.HC 73%

What should AI see? Using the Public's Opinion to Determine the Perception of an AI

Robin Chan, Radin Dardashti, Meike Osinski, Matthias Rottmann, Dominik Brüggemann, Cilia Rücker, Peter Schlicht, Fabian Hüger, Nikol Rummel, Hanno Gottschalk

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

Comments 26 pages, 12 figures

Journal ref AI and Ethics (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.04520 2022-11-30 cs.LG cs.AI 73%

Obtaining Dyadic Fairness by Optimal Transport

Moyi Yang, Junjie Sheng, Xiangfeng Wang, Wenyan Liu, Bo Jin, Jun Wang, Hongyuan Zha

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.11446 2022-01-24 cs.CL cs.AI 73%

Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d'Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, Geoffrey Irving

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments 120 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.11022 2021-06-22 cs.CY cs.AI cs.SY eess.SY 73%

Hard Choices in Artificial Intelligence

Roel Dobbe, Thomas Krendl Gilbert, Yonatan Mintz

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments Pre-print. Shorter versions published at Neurips 2019 Workshop on AI for Social Good and Conference on AI, Ethics and Society 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.04255 2021-02-09 cs.CY cs.AI 73%

AI Development for the Public Interest: From Abstraction Traps to Sociotechnical Risks

McKane Andrus, Sarah Dean, Thomas Krendl Gilbert, Nathan Lambert, Tom Zick

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments 8 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.03021 2017-06-12 cs.AI cs.CY 73%

Ethical Artificial Intelligence - An Open Question

Alice Pavaloiu, Utku Kose

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments 13 pages, 3 figures

Journal ref Journal of Multidisciplinary Developments, 2(2), 2017, 15-27

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02244 2026-08-04 cs.DC cs.SY eess.SY 新提交 71%

Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling

通过资源公平调度实现批量大语言模型(LLM)服务的效率与成本对齐

Dayi Yao, Zijie Zhou

专题命中 AI治理与伦理 :alignment(title)

AI总结 本文针对批量LLM服务中异构请求导致的资源分配低效问题,提出ISJL算法,实现高吞吐量的同时对齐最大驱动批次成本与token计量收入,达到3/4的竞争比下界,平衡了FCFS与LJF的优缺点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09612 2026-05-15 cs.HC 71%

When Thinking Pays Off: Incentive Alignment for Human-AI Collaboration

思考有回报:人机协作中的激励对齐

Joshua Holstein, Patrick Hemmer, Gerhard Satzger, Wei Sun

专题命中 AI治理与伦理 :alignment(title)

AI总结 研究探讨了人机协作中激励结构对过度依赖AI的影响,提出新的激励机制减少过度依赖,通过实验验证其有效性,强调激励设计需与任务和互补性对齐。

Comments Accepted at the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19610 2026-03-24 cs.CV 71%

ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding

ParallelVLM: 无损视频-大语言模型加速与视觉对齐感知并行推测解码

Quan Kong, Yuhao Shen, Yicheng Ji, Huan Li, Cong Wang

机构 * Zhejiang University(浙江大学)

专题命中 AI治理与伦理 :alignment(title)

AI总结 ParallelVLM通过并行化阶段和无偏验证器引导剪枝策略,有效提升视频理解任务的解码效率,实现3.36倍的加速效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23245 2025-10-28 cs.HC cs.MA 71%

Multi-Stakeholder Alignment in LLM-Powered Collaborative AI Systems: A Multi-Agent Framework for Intelligent Tutoring

Alexandre P Uchoa, Carlo E T Oliveira, Claudia L R Motta, Daniel Schneider

专题命中 AI治理与伦理 :alignment(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17253 2025-07-24 cs.RO 71%

Optimizing Delivery Logistics: Enhancing Speed and Safety with Drone Technology

Maharshi Shastri, Ujjval Shrivastav

机构 * Department of Computer Engineering(计算机工程系) L.R. Tiwari College of Engineering(L.R.蒂瓦里工程学院)

专题命中 AI治理与伦理 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09532 2025-02-04 cs.CV 71%

AdaFV: Rethinking of Visual-Language alignment for VLM acceleration

Jiayi Han, Liang Du, Yiwen Wu, Xiangguo Zhou, Hongwei Du, Weibo Zheng

专题命中 AI治理与伦理 :alignment(title)

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15978 2022-11-30 quant-ph 71%

Qubit seriation: Improving data-model alignment using spectral ordering

Atithi Acharya, Manuel Rudolph, Jing Chen, Jacob Miller, Alejandro Perdomo-Ortiz

专题命中 AI治理与伦理 :alignment(title)

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10426 2025-09-29 cs.CY cs.AI cs.HC math.HO 71%

Formalising Human-in-the-Loop: Computational Reductions, Failure Modes, and Legal-Moral Responsibility

Maurice Chiodo, Dennis Müller, Paul Siewert, Jean-Luc Wetherall, Zoya Yasmine, John Burden

机构 * Centre for the Study of Existential Risk(存在风险研究中心) University of Cambridge(剑桥大学) Institute of Mathematics Education(数学教育研究所) University of Cologne(科隆大学) Department of Computer Science and Technology(计算机科学与技术系) DeepFin Research(DeepFin研究) Faculty of Law(法学院) University of Oxford(牛津大学) Leverhulme Centre for the Future of Intelligence(未来智能中心)

专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.AI、cs.CY;trustworthy(comments);AI safety(comments)

Comments 31 pages. Keywords: Human-in-the-loop, Automated decision making system, Human oversight in sociotechnical systems, Oracle machine, AI safety, Trustworthy AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22877 2026-08-11 cs.AI cs.HC cs.RO 版本更新 70%

Physical AI Governance: From Theory to Practice Across Life Cycle

物理人工智能治理:跨越生命周期从理论到实践

Wang Yang, Shaobo Wang, Hongxuan Liu, Xiaoran Cai, Yunyu He, Jingzong Zhou, Mengzhong Ma, Yi Yu, Rohit Sharma, Jingjing Fu, Peng Qi

机构 * Case Western Reserve University(凯斯西储大学) Shanghai Jiao Tong University(上海交通大学) Massachusetts Institute of Technology(麻省理工学院) Columbia University(哥伦比亚大学) University of California, Riverside(加州大学河滨分校) Nanyang Technological University(南洋理工大学) New York University(纽约大学) Salesforce(Salesforce公司) Harvard University(哈佛大学) Stanford University(斯坦福大学)

专题命中 AI治理与伦理 :safety(abstract);trustworthy(abstract);分类 cs.AI

AI总结 本文对物理人工智能治理进行全面综述,综合现有原则形成统一框架,提出五阶段生命周期,通过具体实践展示各阶段治理操作,为构建安全可信且符合社会价值的物理人工智能系统提供参考。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03532 2026-08-05 cs.CL 新提交 70%

Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili

大语言模型中的跨语言偏见:英语与斯瓦希里语的对比分析

Ruolei Zhang, Teddy Njuguna, Yue Feng

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL

AI总结 本研究通过对比GPT-5.2与Gemini 2.5 Flash在英语和斯瓦希里语上的9类人口统计偏见表现,发现偏见会转变而非转移,仅英语偏见审计无法覆盖多语言部署需求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18459 2026-07-22 cs.CY 新提交 70%

Intelligent Cause Prioritisation? An Analysis of AI Policy Priorities and Governance in Africa

智能因果优先级排序?非洲人工智能政策优先级与治理分析

Osaremen Iluobe, Kisso Selvan

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CY

AI总结 该研究聚焦非洲人工智能政策,通过分析相关资料发现,非洲各国政府急于借人工智能加速经济转型,虽重视其机遇,却相对忽视人工智能安全。

Comments 21 pages, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16513 2026-07-21 cs.CY 70%

Competing Visions of Ethical AI: A Case Study of OpenAI

伦理AI的两种视角:OpenAI案例研究

Melissa Wilfley, Mengting Ai, Madelyn Rose Sanfilippo

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CY

AI总结 本研究通过分析OpenAI的公共论述,发现其在伦理AI领域更强调安全性和风险,而非学术或倡导伦理框架。

Comments iConference 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18236 2026-07-07 cs.CY 版本更新 70%

Two Means to an End Goal: Connecting Explainability and Contestability in the Regulation of Public Sector AI

实现目标的两种方式:在公共部门人工智能监管中连接可解释性与可争议性

Timothée Schmude, Mireia Yurrita, Kars Alfrink, Thomas Le Goff, Sebastian Tschiatschek, Tiphaine Viard

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.CY

AI总结 探讨公共部门人工智能监管中可解释性与可争议性,通过对14位专家访谈研究其交叉与实施,明确相关差异及实现难点,从设计角度为人工智能政策提三条建议。

Comments 19 pages main text, 4 figures. Supplementary material is provided

Journal ref Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (pp. 3078-3104). Association for Computing Machinery

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30666 2026-07-01 cs.CY 新提交 70%

Reframing AGI Confrontation with Off Earth Autonomy

重构与地外自主性的AGI对抗

Alexey Potapov

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CY

AI总结 本文质疑AI安全叙事中关于强大智能体必然寻求权力并与人类对抗的结论,提出存在地外自主性路径时,早期合作可取代对抗,并通过决策理论模型分析激励转变及其治理启示。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30661 2026-07-01 cs.CY 新提交 70%

Understanding Censorship in Large Language Models: From Mechanisms to Governance

理解大型语言模型中的审查:从机制到治理

Quanyan Zhu

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CY

AI总结 本文从社会技术视角审视LLM审查,涵盖数据、对齐、政策、推理时审核及法规,分析其表现形式、测量挑战与治理需求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15205 2026-06-23 cs.CY 版本更新 70%

A Peek Behind the Curtain: Using Step-Around Prompt Engineering to Identify Bias and Misinformation in GenAI Models

幕后窥探:使用逐步提示工程识别生成式AI模型中的偏见与虚假信息

Don Hickerson, Mike Perkins

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CY

AI总结 本文探讨逐步提示工程作为对抗性提示技术,用于测试生成式AI的安全护栏和偏见缓解机制,并提出了学术伦理治理框架。

Comments v3 Substantial text changes and addition of structured framework

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00003 2026-06-02 cs.CY cs.CR cs.HC 70%

Learning from Mistakes: Can LLM Self-Recover after Misalignment?

从错误中学习:LLM 在错位后能否自我恢复?

Olga E. Sorokoletova, Francesco Giarrusso, Vincenzo Suriani, Daniele Nardi

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CY

AI总结 研究大语言模型在遭受恶意攻击后,是否具有内在的自我对齐恢复能力,并提出一种建模用户-助手交互安全轨迹并检测恢复趋势的方法。

Comments AAAI'26 Workshop (WS37), Machine Ethics: from formal methods to emergent machine ethics, January 20--27, 2026, Singapore

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27996 2026-06-01 cs.AI 70%

Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure

奖励偏差替代:单轴偏差缓解措施重定向优化压力

Max Lamparth, Daniel Fein, Andreas Haupt, Marcel Hussing, Mykel J. Kochenderfer

机构 * Stanford University(斯坦福大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 AI治理与伦理 :RLHF(abstract,abstract_cn);分类 cs.AI

AI总结 本文提出奖励偏差替代现象,即单轴缓解奖励模型偏差(如减少对长度、谄媚或风格的依赖)会将优化压力转移到相关代理上而非消除,并通过理论证明和实验(如GRPO训练中的长度惩罚导致过度自信)揭示了该问题,建议在评估中纳入策略诱导分布并跟踪多偏差。

Comments Improved readability (mostly appendix D)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27593 2026-05-28 cs.AI cs.MA 70%

Voluntary Collusion with Secret Tools in Competing LLM Agents

竞争性LLM代理中使用秘密工具的合谋行为

Xijie Zeng, Frank Rudzicz

机构 * Dalhousie University(达尔豪斯大学) Vector Institute for Artificial Intelligence(人工智能向量研究所)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本研究通过两个多智能体环境(Liar's Bar和Cleanup)发现,即使工具被明确标注为不公平且有害,大多数LLM代理仍会自愿采用秘密合谋工具以获取战略优势,且仅靠对齐或公平标签无法有效阻止,需明确防护措施。

详情

展开后加载摘要…

URL PDF HTML 收藏