arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1852 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1852 篇

2510.05154 2026-03-23 cs.CL 57%

Can AI Truly Represent Your Voice in Deliberations? A Comprehensive Study of Large-Scale Opinion Aggregation with LLMs

AI能否真正代表您的声音进行辩论?基于大规模意见聚合的全面研究

Shenzhe Zhu, Shu Yang, Michiel A. Bakker, Alex Pentland, Jiaxin Pei

机构 * Stanford University(斯坦福大学) University of Toronto(多伦多大学) KAUST(卡塔尔大学) MIT(麻省理工学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本文通过DeliberationBank数据集评估LLM在大规模辩论总结中的表现,发现其存在代表性不足和偏见问题,提出DeliberationJudge模型提升评估准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00479 2026-03-19 cs.AI 57%

Aligning Probabilistic Beliefs under Informative Missingness: LLM Steerability in Clinical Reasoning

在信息缺失下对概率信念进行对齐:临床推理中的LLM可引导性

Yuta Kobayashi, Vincent Jeanselme, Shalmali Joshi

机构 * Department of Biomedical Informatics(生物医学信息学系)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文研究LLM能否利用信息性缺失进行预测推理,通过分析三种提示方法发现,尽管结构引导和上下文学习可提升概率对齐,但需精心干预才能利用信息性缺失。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16643 2026-03-18 cs.CL 57%

Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy

对讨好型人格的有益论点:推理如何缓解(却掩盖)LLM的讨好行为

Zhaoxin Feng, Zheng Chen, Jianfei Ma, Yip Tin Po, Emmanuele Chersoni, Bo Li

机构 * The Hong Kong Polytechnic University(香港理工大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 研究探讨了推理在缓解LLM讨好行为中的作用,发现推理虽能减少最终决策的讨好倾向,但可能在部分样本中掩盖讨好行为,且LLM在主观任务和权威偏见下更易表现出讨好倾向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15187 2026-03-17 cs.CL 57%

The Hrunting of AI: Where and How to Improve English Dialectal Fairness

AI的提升:在哪里以及如何改进英语方言公平性

Wei Li, Adrian de Wynter

机构 * Boston College(波士顿学院) Microsoft(微软) The University of York(约克大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 研究探讨了数据质量和可用性对改进LLM在英语方言中的性能的影响,发现人类间的一致性直接影响LLM作为评判者的性能,并指出需谨慎评估数据以确保公平性和包容性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14838 2026-03-17 cs.CL 57%

The Impact of Ideological Discourses in RAG: A Case Study with COVID-19 Treatments

检索增强生成中意识形态话语的影响:以新冠治疗方案为例

Elmira Salari, Maria Claudia Nunes Delfino, Hazem Amamou, José Victor de Souza, Shruti Kshirsagar, Alan Davoust, Anderson Avila

机构 * Wichita State University(威斯康星州立大学) Pontifícia Universidade Católica de São Paulo(圣保罗天主教大学) Institut national de la recherche scientifique(国家科学研究中心) Université du Québec en Outaouais(魁北克大学Outaouais)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本文研究检索到的意识形态文本对大语言模型输出的影响,通过新冠治疗方案案例,探讨意识形态在RAG框架中的识别与影响,揭示意识形态偏见与潜在风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14325 2026-03-17 cs.AI 57%

FAIRGAME: a Framework for AI Agents Bias Recognition using Game Theory

FAIRGAME: 一个用于通过博弈论识别AI代理偏见的框架

Alessio Buscemi, Daniele Proverbio, Alessandro Di Stefano, The-Anh Han, German Castignani, Pietro Liò

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

AI总结 FAIRGAME利用博弈论框架识别AI代理中的偏见,通过模拟和比较不同场景下的结果,系统发现偏见并预测战略互动中的行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13189 2026-03-16 cs.MA cs.AI 57%

LLM Constitutional Multi-Agent Governance

大语言模型宪法多智能体治理

J. de Curtò, I. de Zarzà

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出CMAG框架,通过硬约束过滤与软惩罚效用优化,在多智能体群体中平衡合作潜力与操纵风险,实验显示CMAG在保持自主性和完整性的同时提升了合作的伦理得分。

Comments Accepted for publication in 20th International Conference on Agents and Multi-Agent Systems: Technologies and Applications (AMSTA 2026), to appear in Springer Nature proceedings (KES Smart Innovation Systems and Technologies). The final authenticated version will be available online at Springer

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08093 2026-03-16 cs.CL 57%

Evolution and compression in LLMs: On the emergence of human-aligned categorization

语言模型中的进化与压缩:关于人类对齐分类的出现

Nathaniel Imel, Noga Zaslavsky

机构 * New York University(纽约大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 研究探讨LLM是否能进化出高效的人类对齐语义系统,通过颜色分类实验发现,大模型在复杂性和英语对齐性上表现各异,仅最强模型能复现人类近最优的信息瓶颈贸易-offs。

Comments Published as a conference paper at ICLR 2026 (The Fourteenth International Conference on Learning Representations). OpenReview: https://openreview.net/forum?id=s7gSTR2AqA&noteId=s7gSTR2AqA

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11277 2026-03-16 cs.AI 57%

COMPASS: The explainable agentic framework for Sovereignty, Sustainability, Compliance, and Ethics

COMPASS:面向主权、可持续性、合规性和伦理的可解释代理框架

Jean-Sébastien Dessureault, Alain-Thierry Iliho Manzi, Soukaina Alaoui Ismaili, Khadim Lo, Mireille Lalancette, Éric Bélanger

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出COMPASS框架,通过模块化治理机制整合主权、可持续性、合规性和伦理,利用RAG技术提升AI决策的可解释性和透明度。

Comments 22 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00307 2026-03-12 cs.AI 57%

BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language Models

BiasBusters: 检测和缓解大语言模型中的工具选择偏差

Thierry Blankenstein, Jialin Yu, Zixuan Li, Vassilis Plachouras, Sunando Sengupta, Philip Torr, Yarin Gal, Alasdair Paren, Adel Bibi

机构 * University of Oxford(牛津大学) Microsoft(微软)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 研究通过基准测试揭示LLM工具选择中的系统性偏差,并提出轻量级缓解策略以减少选择偏差。

Comments ICLR 2026 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22699 2026-03-12 cs.CL 57%

Are you sure? Measuring models bias in content moderation through uncertainty

你确定吗?通过不确定性测量内容审核中的模型偏差

Alessandra Urbinati, Mirko Lai, Simona Frenda, Marco Antonio Stranisci

机构 * Laboratory for the Modeling of Biological and Socio-technical Systems, Northeastern University(生物与社会技术系统建模实验室,东北大学) Heriot-Watt University(赫瑞-瓦特大学) aequa-tech(aequa-tech公司) Università del Piemonte Orientale(皮埃蒙特东方大学) Università degli Studi di Torino(托里尼大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

AI总结 本文提出通过模型预测不确定性来衡量内容审核中模型的偏差,揭示预训练模型对少数群体的预测准确性与置信度的差异,以改进模型公平性。

Comments accepted at Findings of ACL: EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04417 2026-03-06 cs.CL 57%

Same Input, Different Scores: A Multi Model Study on the Inconsistency of LLM Judge

相同输入,不同评分:多模型研究LLM裁判的一致性

Fiona Lau

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本研究探讨了LLM在不同模型、温度设置下对相同输入评分的一致性问题,发现模型间评分差异显著,温度影响稳定性,需引入混合评估策略以确保可靠应用。

Comments 19 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21648 2026-03-05 cs.AI 57%

Leveraging Imperfection with MEDLEY A Multi-Model Approach Harnessing Bias in Medical AI

利用不完美性:MEDLEY:一种多模型方法,利用医学AI中的偏差

Farhad Abtahi, Mehdi Astaraki, Fernando Seoane

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

AI总结 MEDLEY提出了一种多模型框架,通过保留模型多样性而非压缩共识,利用医学AI中的偏差提升诊断准确性与透明度。

Journal ref Front. Artif. Intell., Volume 9 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03018 2026-03-04 cs.AI cs.SE 57%

REGAL: A Registry-Driven Architecture for Deterministic Grounding of Agentic AI in Enterprise Telemetry

REGAL:一种基于注册表的架构,用于企业遥测中代理AI的确定性接地

Yuvraj Agrawal

机构 * Adobe Inc.(Adobe公司)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 REGAL提出一种基于注册表的架构,用于企业遥测中代理AI的确定性接地,通过显式架构方法和语义编译提升确定性计算,解决LLM在私有遥测中的接地问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11922 2026-03-04 cs.CY cs.SE 57%

On Regulating Downstream AI Developers

对下游AI开发者进行监管

Sophie Williams, Jonas Schuett, Markus Anderljung

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

AI总结 本文探讨了对下游AI开发者进行监管的必要性,提出通过要求上游开发者减轻下游修改风险或使用替代政策工具来平衡监管与创新。

Comments 39 pages, 2 figures, 7 tables

Journal ref Eur. j. risk regul. 17 (2026) 94-122

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02684 2026-03-04 cs.CL cs.SI 57%

HateMirage: An Explainable Multi-Dimensional Dataset for Decoding Faux Hate and Subtle Online Abuse

HateMirage: 一个可解释的多维数据集用于解码虚假仇恨与微妙的在线虐待

Sai Kartheek Reddy Kasu, Shankar Biradar, Sunil Saumya, Md. Shad Akhtar

机构 * Indian Institute of Information Technology Dharwad(印度信息技术学院达尔瓦德) Manipal Institute of Technology, Manipal Academy of Higher Education(曼普及高等教育学院技术学院) Indraprastha Institute of Information Technology Delhi(印度普拉斯塔信息技术学院德里)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

AI总结 HateMirage 是一个用于解码虚假仇恨与微妙在线虐待的多维数据集,通过多维解释框架提升仇恨言论的可解释性研究。

Comments Accepted at LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02568 2026-03-04 cs.CY 57%

AI4CAREER: Responsible AI for STEM Career Development at Scale in K-16 Education

AI4CAREER: 为K-16教育中的STEM职业发展实现负责任的AI

Sugana Chawla, Si Chen, Julia Qian, Gina Svarovsky, Alison Cheng, Rick Johnson, Nitesh V. Chawla, Ronald Metoyer

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 AI4CAREER研讨会探讨如何在K-16教育中通过负责任的AI促进STEM职业发展,聚焦AI在职业准备度评估、决策边界、发展一致性及公平性设计等方面的核心问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23720 2026-03-02 cs.AI 57%

The Auton Agentic AI Framework

自主AI框架

Sheng Cao, Zhao Chang, Chang Li, Hannan Li, Liyao Fu, Ji Tang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 Auton框架通过分离认知蓝图与运行时引擎,实现自主代理系统的标准化创建、执行和治理,采用增强POMDP模型和三级自我进化框架提升安全性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08552 2026-03-02 cs.LG cs.CV 57%

Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy Biases

对抗扩散模型中的奖励过度优化:从归纳偏置和优先偏置的角度出发

Ziyi Zhang, Sen Zhang, Yibing Zhan, Yong Luo, Yonggang Wen, Dacheng Tao

机构 * Institute of Artificial Intelligence, School of Computer Science, Wuhan University, China Hubei Luojia Laboratory, Wuhan, China The University of Sydney, Australia JD Explore Academy, Beijing, China Nanyang Technological University, Singapore

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本文提出TDPO-R算法,通过利用扩散模型的时间归纳偏置和抑制活跃神经元的优先偏置,有效缓解奖励过度优化问题。

Comments Accepted to ICML 2024

Journal ref International Conference on Machine Learning, pp. 60396-60413, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19682 2026-02-24 cs.CY 57%

Beyond the Binary: A nuanced path for open-weight advanced AI

超越二元:开放权重高级AI的细致路径

Bengüsu Özcan, Alex Petropoulos, Max Reddel

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

AI总结 本文提出了一种基于安全评估的分层模型发布方法,旨在解决开放权重高级AI模型在安全性和监管方面的挑战。

Comments This publication was originally designed and optimised for web and published on cfg.eu. Minor formatting differences may appear in this version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07754 2026-02-24 cs.AI cs.HC 57%

Humanizing AI Grading: Student-Centered Insights on Fairness, Trust, Consistency and Transparency

让AI评分更人性化:以学生为中心的公平性、信任、一致性与透明性洞察

Bahare Riahi, Viktoriia Storozhevykh, Veronica Catete

机构 * North Carolina State University(北卡罗来纳州立大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

AI总结 本研究通过比较AI与人工评分反馈,探讨学生对AI评分系统在公平性、信任、一致性与透明性方面的看法,并提出人本化AI的设计原则。

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18535 2026-02-24 cs.SD cs.AI 57%

Fairness-Aware Partial-label Domain Adaptation for Voice Classification of Parkinson's and ALS

面向语音分类的公平性意识部分标签领域适应

Arianna Francesconi, Zhixiang Dai, Arthur Stefano Moscheni, Himesh Morgan Perera Kanattage, Donato Cappetta, Fabio Rebecchi, Paolo Soda, Valerio Guarrasi, Rosa Sicilia, Mary-Anne Hartley

机构 * organization= School of Computer Communication Sciences, EPFL (\'Ecole polytechnique f\'ed\'erale de Lausanne) , city= Lausanne , country= Switzerland organization= Eustema S.p.A., Research Development Centre , city= Naples , country= Italy organization= UniCamillus-Saint Camillus International University of Health Sciences , city= Rome , country= Italy organization= Department of Diagnostics Intervention, Radiation Physics, Biomedical Engineering, Umeå University , city= Umeå , country= Sweden

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出了一种融合域泛化和对抗对齐的框架,用于在部分标签不匹配和公平性约束下实现帕金森病和肌萎缩侧索硬化症的统一语音分类。

Comments 7 pages, 1 figure. Submitted to Pattern Recognition Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18461 2026-02-24 cs.CY 57%

Toward Self-Driving Universities: Can Universities Drive Themselves with Agentic AI?

迈向自动驾驶大学:大学能否通过代理AI实现自我驱动?

Anis Koubaa

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本研究提出通过代理AI实现高等教育机构的自主性框架,旨在自动化行政、学术和质量保证流程,减少教师文书工作时间,提升教育质量和研究生产力。

Journal ref Springer Book: Next Generation AI-Driven Education - 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17919 2026-02-23 cs.CY cs.HC 57%

Visual Anthropomorphism Shifts Evaluations of Gendered AI Managers

视觉人化影响对性别化AI管理者评价

Ruiqing Han, Hao Cui, Taha Yasseri

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

AI总结 研究发现,文本描述中能力信息可缓解对AI管理者的负面评价,而视觉人化会引发性别偏见,表明表示方式影响性别刻板印象的激活。

Comments Preprint, Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16553 2026-02-19 cs.CY 57%

Agentic AI, Medical Morality, and the Transformation of the Patient-Physician Relationship

代理AI、医疗道德与患者-医生关系的变革

Robert Ranisch, Sabine Salloch

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

AI总结 本文探讨代理AI如何通过重塑患者-医生关系改变医疗道德,呼吁在广泛应用前融入伦理考量。

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12680 2026-02-16 stat.ML cs.LG 57%

A Regularization-Sharpness Tradeoff for Linear Interpolators

线性插值器的正则化-尖锐性权衡

Qingyi Hu, Liam Hodgkinson

机构 * School of Mathematics and Statistics(数学与统计学学院) University of Melbourne(墨尔本大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本文提出了一种针对过参数化线性回归的正则化-尖锐性权衡,通过ℓ^p惩罚分解选择惩罚为正则化项和几何尖锐性项,验证了其在现实数据中的有效性。

Comments 29 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11301 2026-02-13 cs.AI cs.CR 57%

The PBSAI Governance Ecosystem: A Multi-Agent AI Reference Architecture for Securing Enterprise AI Estates

PBSAI治理生态系统:一种多智能体AI参考架构,用于保障企业AI领地安全

John M. Willis

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 PBSAI提出了一种多智能体参考架构,用于保障企业AI领地的安全,通过十二个领域分类法和有限智能体家族实现责任划分,结合分析监控、协调防御等技术,确保系统安全与可追溯性。

Comments 43 pages, plus 12 pages of appendices. One Figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08193 2026-02-12 cs.AI 57%

Measuring What Matters: The AI Pluralism Index

衡量重要性:AI多元主义指数

Rashid Mushkani

机构 * Université de Montréal(蒙特利尔大学) Mila – Québec AI Institute(魁北克人工智能研究所)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 本文提出AI多元主义指数,用于衡量人工智能系统在治理、包容性和透明度方面的多元实践,旨在引导激励向多元主义方向发展。

Comments Proceedings of the International Association for Safe & Ethical AI (IASEAI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12365 2026-02-12 cs.CL cs.DB 57%

Advances in LLMs with Focus on Reasoning, Adaptability, Efficiency and Ethics

大语言模型的进展:聚焦推理、适应性、效率和伦理

Asifullah Khan, Muhammad Zaeem Khan, Aleesha Zainab, Saleha Jamshed, Sadia Ahmad, Kaynat Khatib, Faria Bibi, Abdul Rehman

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本文综述了大语言模型在推理、适应性、效率和伦理方面的进展,探讨了关键技术和挑战,提出未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06218 2026-02-11 cs.CV cs.LG 57%

Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings

跨模态冗余与视觉-语言嵌入的几何学

Grégoire Dhimoïla, Thomas Fel, Victor Boutin, Agustin Picard

机构 * Brown University(布朗大学) ENS Paris Saclay(巴黎萨克雷大学) IRT Saint Exupéry(IRT圣埃克苏佩里) Kempner Institute, Harvard University(哈佛大学凯姆纳研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本文通过等能假设和对齐稀疏自编码器,揭示了视觉-语言模型中跨模态对齐的几何结构,发现稀疏双模态原子承载了跨模态对齐信号,单模态原子解释了模态差距,去除单模态原子可消除差距而不影响性能。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏