arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2605.01451 2026-05-05 cs.CL 61%

Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language models

对基于AI的紧急警务调度中的种族偏见进行审计:对十一种大型语言模型的跨语言评估

William Guey, Wei Zhang, Pierrick Bougault, Yi Wang, Bertan Ucar, Vitor D. de Moura, José O. Gomes

机构 * Department of Industrial Engineering, Tsinghua University(清华大学工业工程系) School of Social Sciences, Tsinghua University(清华大学社会科学部) Department of Industrial Engineering, Federal University of Rio de Janeiro(里约热内卢联邦大学工业工程系)

专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.CL

AI总结 本文通过跨语言框架评估11种模型,在19800个输出中发现当事件严重性模糊时种族偏见系统性出现,但当操作优先级由通话内容确定时偏见消失。偏见程度因种族轴而异,宗教外观影响最大,性别次之,种族最小。语言间偏见转移不一致,性别偏见在中文中放大,种族偏见在英文中更明显。

Comments 26 pages, 7 figures. Submitted to Humanities and Social Sciences Communications (Nature) collection on Artificial Intelligence and Emerging Technologies in Public Safety. Code and data: https://github.com/williamguey/llmdispatchbias

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11839 2026-04-15 cs.CR cs.AI 61%

Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents

超越静态沙箱:为自主AI代理的学得能力治理

Bronislav Sidik, Lior Rokach

机构 * Institute for Applied AI Research(应用人工智能研究所) Faculty of Computer and Information Science(计算机与信息科学学院) Ben-Gurion University of the Negev(贝内-约尔大学)

专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.AI

AI总结 本文提出Aethelgard框架,通过学得策略实现AI代理的最小必要能力集,解决能力过度配置问题。

Comments 17 pages (9 content pages), 2 figures, 7 tables. Submitted to NeurIPS 2026 Agent Safety Workshop. Code and dataset available at https://github.com/sidikbro/aethelgard

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22037 2025-12-05 cs.CY 61%

What AI Speaks for Your Community: Polling AI Agents for Public Opinion on Data Center Projects

人工智能为你的社区发声:通过AI代理收集数据中心项目公众意见

Zhifeng Wu, Yuelin Han, Shaolei Ren

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY;trustworthy(comments)

AI总结 本文提出AI代理调查框架,利用大型语言模型评估社区对数据中心项目的意见,以指导负责任的AI发展。

Comments 35 Pages. Accepted to NeurIPS 2025 Workshop on Socially Responsible and Trustworthy Foundation Models (ResponsibleFM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.09222 2025-08-22 cs.CY 61%

Democratic AI is Possible. The Democracy Levels Framework Shows How It Might Work

Aviv Ovadya, Kyle Redman, Luke Thorburn, Quan Ze Chen, Oliver Smith, Flynn Devine, Andrew Konya, Smitha Milli, Manon Revel, K. J. Kevin Feng, Amy X. Zhang, Bilva Chandra, Michiel A. Bakker, Atoosa Kasirzadeh

专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CY

Comments 31 pages. Accepted to the position paper track at ICML 2025. A previous version was presented at the Pluralistic Alignment Workshop at NeurIPS 2024. For ongoing work, see: https://democracylevels.org

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.05617 2025-05-20 cs.LG cs.AI cs.CL cs.CV cs.CY 61%

Debiasing Methods for Fairer Neural Models in Vision and Language Research: A Survey

Otávio Parraga, Martin D. More, Christian M. Oliveira, Nathan S. Gavenski, Lucas S. Kupssinskü, Adilson Medronha, Luis V. Moura, Gabriel S. Simões, Rodrigo C. Barros

机构 * Machine Learning Theory and Applications (MALTA) Lab, PUCRS(机器学习理论与应用(MALTA)实验室,PUCRS)

专题命中 AI治理与伦理 :分类 cs.CL、cs.AI、cs.CY;trustworthy(comments)

Comments Submitted to ACM Computing Surveys - Special Issue on Trustworthy AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11579 2025-01-22 cs.CL 61%

HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection

Theo King, Zekun Wu, Adriano Koshiyama, Emre Kazim, Philip Treleaven

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL;safety(comments)

Comments NeurIPS 2024 SoLaR Workshop and NeurIPS 2024 Safety Gen AI Workshop

Journal ref NeurIPS Safe Generative AI Workshop 2024; Workshop on Socially Responsible Language Modelling Research 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.09362 2024-10-02 cs.LG 61%

Long-Term Fairness with Unknown Dynamics

Tongxin Yin, Reilly Raab, Mingyan Liu, Yang Liu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG;trustworthy(comments)

Comments Best paper runner-up at ICLR 2023 Workshop on Trustworthy and Reliable Large-Scale Machine Learning Models (Non Archival)

Journal ref Advances in Neural Information Processing Systems 36 (NeurIPS 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09209 2023-07-19 cs.CL cs.AI cs.CY cs.LG 61%

Automated Ableism: An Exploration of Explicit Disability Biases in Sentiment and Toxicity Analysis Models

Pranav Narayanan Venkit, Mukund Srinath, Shomir Wilson

专题命中 AI治理与伦理 :分类 cs.CL、cs.AI、cs.CY;trustworthy(journal_ref)

Comments TrustNLP at ACL 2023

Journal ref Proceedings at The Third Workshop on Trustworthy Natural Language Processing collocated at the 61st Annual Meeting Of The Association For Computational Linguistics. 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.00168 2022-11-02 cs.CV cs.LG 61%

Improving Fairness in Image Classification via Sketching

Ruichen Yao, Ziteng Cui, Xiaoxiao Li, Lin Gu

专题命中 AI治理与伦理 :trustworthy(abstract,comments);分类 cs.LG

Comments 8 pages, 2 figures. To appear in 2022 Trustworthy and Socially Responsible Machine Learning (TSRML 2022) co-located with NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21363 2026-08-25 cs.AI cs.CR 新提交 57%

AIREP: A Protocol for Per-Decision Evidence in AI Runtime Governance

AIREP:AI运行时治理中逐决策证据的协议

Ali Toygar Abak

机构 * Phionyx Research(菲奥尼克斯研究院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 AIREP是一种用于记录自动化AI运行时治理决策的协议,以签名对象形式存储决策,通过SHA-256哈希链防篡改,可供相关AI运行时采用。

Comments 8 pages. Reference implementation and two-verifier conformance kit: https://github.com/halvrenofviryel/ai-runtime-evidence-protocol

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18774 2026-08-20 cs.CV cs.LG 新提交 57%

MIFR: A Modality-Invariant and Fair Representation Framework for Skin Disease Classification

MIFR:用于皮肤病分类的模态不变公平表示框架

Asonyu Senge Njih, Yvan Guifo Fodjo, Vianney Kengne Tchendji, Jerry Lacmou Zeutouo, Kerol Djoumessi

机构 * University of Dschang(德昌大学) Université Paris-Panthéon-Assas(巴黎先贤祠-阿萨斯大学) Université de Picardie Jules Verne(儒勒·凡尔纳皮卡第大学) Hertie Institute for AI in Brain Health, University of Tübingen(蒂宾根大学赫蒂脑健康人工智能研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本研究提出MIFR框架,通过配对临床与皮肤镜图像的多目标损失训练,实现皮肤病分类的模态不变性与公平性,在多数据集上验证了其预测性能与公平性。

Comments Accepted for publication at the First Workshop on Advancing African Medical AI through Global Integration (AFRICAI)-MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06412 2026-08-20 cs.CY 版本更新 57%

Brokerage in the Black Box: Swing States, Strategic Ambiguity, and the Global Politics of AI Governance

黑箱中的中介作用: swing州、战略模糊性与人工智能治理的全球政治

Ha-Chi Tran

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本研究探讨了技术swing州如何通过战略模糊性和制度透明性互动,调解大国技术竞争,影响全球人工智能治理的制度设计与政策结果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17516 2026-08-19 cs.CL 新提交 57%

Effects of Answer Format Variation on Gender Bias in Large Language Models

答案格式变化对大型语言模型中性别偏见的影响

Ksenia Merzlyakova, Sebastian Padó, Franziska Weeber

机构 * Institute for Natural Language Processing, University of Stuttgart(斯图加特大学自然语言处理研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 该研究探讨答案格式变化对LLMs性别偏见测量的影响,通过评估三种指令微调模型在不同格式下的表现,发现格式会显著改变测量结果,强调需将答案格式纳入LLM评估。

Comments 6th Workshop on Computational Linguistics for the Political and Social Sciences (CPSS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16394 2026-08-18 cs.AI cs.IR 新提交 57%

Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152

在块内思考:使用大语言模型生成符合法规的场景的RegulaRAG——以联合国第152号法规为例

Vahid Zolfaghari, Nenad Petrovic, AndrÉ Schamschurko, Alois Knoll

机构 * Technical University of Munich(慕尼黑工业大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 针对LLMs难以结合冗长分层标准的问题,提出RegulaRAG流水线,经实验其在UN R152数据集上元分数最高且鲁棒性强,优于基线系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15354 2026-08-18 cs.AI 新提交 57%

Incoherent by Design? On the Moral Self-Consistency of LLMs

天生不连贯?大型语言模型的道德自我一致性研究

Pegah Nokhiz, Aravinda Kanchana Ruwanpathirana, Helen Nissenbaum

机构 * Cornell University(康奈尔大学) Cornell Tech(康奈尔科技学院) Nanyang Technological University(南洋理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本研究针对GPT、Mistral、Llama等LLM,在义务论等三大伦理框架下,发现其道德推理存在最高78%的矛盾率,内部不连贯是AI对齐的必要前提。

Comments 88 pages; pages 16 to 88 are the appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12669 2026-08-14 cs.CY 新提交 57%

From Fair Representation to Just Recognition in Generative AI

从生成式AI中的公平表征到公正承认

Severin Engelmann, Daniel Susser

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 该研究针对生成式AI的公平挑战,指出现有表征公平策略存在局限,提出借鉴参与平等理论,从表征公平转向承认正义以解决相关问题。

Comments Accepted for publication at the 2026 AAAI/ACM Conference on AI, Ethics, and Society (AIES)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11207 2026-08-13 cs.AI 新提交 57%

Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes

用于实现协作式对话结果的多大型语言模型智能体系统的动态管控

Alexander Liss, Nicholas Desmond, Santiago Gil Gallego

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文针对多LLM智能体对话崩溃问题,提出体验编排器EO管控层,通过三种机制提升顾问联系率,在6万次模拟中取得显著效果,为多智能体协作提供新方案。

Comments 13 pages, 3 figures, 3 tables. Submitted to AI Engineer World's Fair 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15770 2026-08-12 cs.CV cs.LG 版本更新 57%

Concept Labels Are Not Enough: Rethinking Concept Bottleneck Models through Representation Integrity

通过解耦概念瓶颈缓解多媒体识别中的虚假背景偏差

Gaoxiang Huang, Songning Lai, Yutao Yue

机构 * HKUST(GZ)(香港科技大学(广州))

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本文提出轻量解耦概念瓶颈模型LDCBM,通过滤波分组损失和联合概念监督提升视觉模式与概念的对齐,实验证明其在概念和分类准确率上优于现有CBMs,且计算复杂度更低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09857 2026-08-11 cs.RO cs.AI 新提交 57%

Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy

智能体 harness:面向机器人自主的大语言模型驱动验证层

Rohan Bhagra, Mahantesh Halapannavar, Uddhav Bhattarai

机构 * Carnegie Mellon University(卡内基梅隆大学) Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 针对机器人规划模型的安全与伦理风险,提出LLM驱动的验证层作为中间件管控计划,实现近85%的类别准确率、97%的对抗性攻击遏制率,为机器人自主提供可靠保障。

Comments 7 pages. Not yet finalized for conference submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08284 2026-08-11 cs.AI 新提交 57%

Fair on the Surface? Benchmarking Hidden-Output Fairness Gaps in LLM Recommenders

表面上公平?对LLM推荐器的隐藏输出公平性差距进行基准测试

Chan Aristella Lu, Arya Fayyazi, Junhao Zhang, Saeid Shokoufa, Yue Xing, Zhen Xiang, Kyu Hyung Lee, Mehdi Kamal, Massoud Pedram

机构 * University of Georgia(佐治亚大学) University of Southern California(南加州大学) Carnegie Mellon University(卡内基梅隆大学) Michigan State University(密歇根州立大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 该研究提出首个联合评估LLM推荐器可观测输出与隐藏表示公平性的基准FairGap,发现二者存在根本张力,现有框架无法诊断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08220 2026-08-11 cs.AI 新提交 57%

Metanormative Theory for RL-Based Moral Agents

基于强化学习的道德智能体的元规范理论

Aleks Knoks, Marija Slavkovik

机构 * University of Luxembourg(卢森堡大学) University of Bergen(卑尔根大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文针对强化学习(RL)道德智能体设计中哲学文献被边缘化的问题,提炼元规范理论的相关理念审视RL架构,以明确RL智能体道德行为的判定标准,为评估相关RL方法奠定基础。

Comments The 14th International Workshop on Engineering Multi-Agent Systems (EMAS 2026) held May 25-26, 2026 Co-located with AAMAS 2026 Paphos, Cyprus

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07515 2026-08-11 cs.CY 新提交 57%

Bridging AI Risk Frameworks: Reconciling ISO/IEC 42001, the NIST AI Risk Management Framework, and the EU AI Act into a Uni ed Governance Taxonomy

衔接AI风险框架:将ISO/IEC 42001、美国国家标准与技术研究院AI风险管理框架(NIST AI RMF 1.0)及欧盟AI法案整合为统一治理分类体系

Vinod Dhiman

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY

AI总结 本文整合ISO/IEC 42001、NIST AI RMF和欧盟AI法案,构建统一AI治理分类体系,提出实施模型与示例,助力组织同时满足三者要求且无需重复付出。

Comments 17 pages, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07477 2026-08-11 cs.HC cs.AI 新提交 57%

Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human Resources

面向伦理人工智能的设计:用于人力资源的自动机器学习(AutoML)中提升公平性与用户体验的人机交互(HCI)功能考量

Sundaraparipurnan Narayanan

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 该论文结合多视角研究AutoML在人力资源招聘中的公平性,发现现有平台存在缺陷,提出五维度HCI公平评估框架,建议将公平性嵌入AutoML设计以提升伦理可持续性。

Comments 295 pages; Doctoral Thesis;28 figures; 42 tables; 248 references

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20255 2026-08-10 cs.CR cs.AI 版本更新 57%

The Ethics of Autonomous AI Agents for Offensive Security

用于进攻性安全的自主人工智能代理的伦理问题

Andreas Happe, Jürgen Cito, Jasmin Wachter

机构 * TU Wien(维也纳技术大学) University of Klagenfurt(克雷格福大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 研究使用自主人工智能代理进行进攻性安全时的伦理问题,分析道德归因在多方的扩散及技术对利益相关者的影响,通过研究其不确定性等特性及攻防成本不对称性,为现有框架不适用于此情况提供分层建议。

Comments accepted at FAIEMA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06353 2026-08-07 cs.GT cs.AI cs.MA 新提交 57%

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

资源权威:部署AI代理参与式治理的机制设计模型

Praphul Chandra, Sujit Gujar, Ganesh Ghalme

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 本研究提出资源权威机制设计模型,用于通过计算预算使授权自我执行,实现对已部署AI代理的参与式治理,同时指出被治理代理操纵治理 electorate 是核心开放问题。

Comments 22 pages, 9 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06246 2026-08-07 cs.LG 新提交 57%

A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance

用于人工智能治理的后训练适应技术的六维分类法

Fardin Afdideh, Fernando Seoane, Farhad Abtahi

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本综述构建了后训练适应技术的六维分类法,梳理技术间关系,为AI治理提供术语支持,并指出该领域的开放挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06115 2026-08-07 cs.AI 新提交 57%

Mind the Gaps: Mixture-of-Minds for Human Simulation

关注差距:用于人类模拟的思维混合模型

Pranav Dahiya

机构 * Semilattice(半格)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出Anacreon受众模拟模型,基于Gemma 4 12B构建思维混合模型,通过聚类个体、情感链等技术,在个体层面序数对齐达0.775,缩小了群体与个体模拟的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04382 2026-08-06 cs.LG stat.ML 新提交 57%

Non-asymptotic implicit bias of logistic regression at early-stage gradient descent dynamics

早期梯度下降动力学下逻辑回归的非渐近隐式偏置

Han Bao

机构 * The Institute of Statistical Mathematics(统计数理研究所) Tohoku University(东北大学) RIKEN AIP(理化学研究所人工智能项目)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本研究揭示早期梯度下降动力学中逻辑回归参数向量与最大间隔方向弱对齐的机制,给出对齐迭代次数的紧上界,为理解训练时长与泛化性能的关联提供理论支撑。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03670 2026-08-05 cs.CY 新提交 57%

Accountability Asymmetry and Structural Trust in Autonomous AI Systems

自主AI系统中的问责不对称性与结构性信任

Nathan DeBardeleben

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本文针对自主AI系统中问责不对称导致的信任问题,提出将其治理视为基础设施可靠性问题,通过工程异质性方案,即独立监测审查补充流程,解决该问题。

Comments 16 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02685 2026-08-05 cs.SE cs.AI 新提交 57%

BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests

BulkPR-Bench:针对交互拉取请求的队列级治理基准

Zetong Xiong, Qiao Zhao, Jun Zhang, Xueying Lyu, Zhi Li, Yixiang Tu, Xiaowen Yang, Yunjie Zhang, Yufeng Wang, Zhe Zhang, Kaize Yu, Hanwen Du, Zhongkai Sun, Zhuoxin Liu, Zekun Lin, Jianwen Yang, Ruining Chen, Ying Zhang, Tingxuan Pan, Ke Chen, Shubin Han, Chuanhao Sun, Yehua Yang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 该研究推出BulkPR-Bench基准,针对交互PR队列治理,实验显示现有模型在该任务上表现优于顺序基准,但仍存在较大提升空间。

Comments 12 pages, 5 figures. Artifact: https://github.com/Eureka246/BulkPR-Bench-Release ; archived artifact: https://doi.org/10.5281/zenodo.21717780

详情

展开后加载摘要…

URL PDF HTML 收藏