arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2305.02748 2026-02-09 cs.AI cs.CY cs.MA 62%

A computational framework for human values

人类价值观的计算框架

Nardine Osman, Mark d'Inverno

机构 * Artificial Intelligence Research Institute (IIIA-CSIC)(人工智能研究 institute (IIIA-CSIC))

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出一个基于社会科学的正式框架,用于系统研究如何通过人类价值观设计伦理人工智能。

Journal ref Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2024), pp. 1531-1539

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12300 2026-02-05 cs.CY cs.AI 62%

Mutually Assured Deregulation

相互保障的去监管

Gilad Abiri

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 本文指出,去监管反而会加剧安全风险,强调需加强监管以应对人工智能带来的国家安全挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25471 2026-02-02 cs.AI cs.CY 62%

An Aristotelian ontology of instrumental goals: Structural features to be managed and not failures to be eliminated

一种亚里士多德式的工具性目标本体论:需管理而非消除的结构特征

Willem Fourie

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出一种亚里士多德式的工具性目标本体论,将工具性目标视为高级AI系统需管理的特征,而非需消除的异常。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20449 2026-01-29 cs.LG cs.AI 62%

Fair Recourse for All: Ensuring Individual and Group Fairness in Counterfactual Explanations

公平的 recourse 为所有:在反事实解释中确保个体和群体公平性

Fatima Ezzeddine, Obaida Ammar, Silvia Giordano, Omran Ayoub

机构 * University of Applied Sciences and Arts of Southern Switzerland(瑞士南方应用科学与艺术大学) Università della Svizzera italiana(瑞士意大利大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种基于强化学习的方法,生成同时满足个体和群体公平性的反事实解释,以提升XAI的公平性和透明度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16237 2026-01-26 cs.MA cs.AI cs.CY cs.SE 62%

Computational Foundations for Strategic Coopetition: Formalizing Collective Action and Loyalty

战略竞合的计算基础:形式化集体行动与忠诚

Vik Pant, Eric Yu

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出了一种基于忠诚度调节的效用函数,用于分析竞合环境中集体行动问题的产生与解决,通过实验验证展示了忠诚度对努力差异的显著影响。

Comments 68 pages, 22 figures. Third technical report in research program; should be read with companion arXiv:2510.18802 and arXiv:2510.24909. Adapts and extends complex actor material from Pant (2021) doctoral dissertation, University of Toronto

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11893 2026-01-22 cs.CY cs.AI 62%

Beyond Automation: Rethinking Work, Creativity, and Governance in the Age of Generative AI

超越自动化:在生成式人工智能时代重新思考工作、创造力与治理

Haocheng Lin

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文探讨生成式AI对工作、创造力和治理的影响,提出包容性AI治理框架,强调UBI与技能发展等的结合以应对AI带来的挑战。

Comments Improved structure and clarity of the introduction and literature review; explicit articulation of the paper's contributions; refined the integration of AI across labour, UBI, and governance

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11583 2026-01-21 cs.CY cs.AI cs.MA 62%

Bit-politeia: An AI Agent Community in Blockchain

Bit-politeia: 区块链上的一个AI代理社区

Xing Yang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 Bit-politeia通过区块链和AI代理构建公平高效的资源分配系统,以减少传统同行评审中的偏见和资源集中问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09114 2026-01-14 cs.AI cs.CY 62%

AI TIPS 2.0: A Comprehensive Framework for Operationalizing AI Governance

AI TIPS 2.0: 一个全面的AI治理实施框架

Pamela Gupta

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

AI总结 AI TIPS 2.0是一个全面的AI治理实施框架,旨在解决AI部署中的三个关键治理挑战,通过提供定制化的治理方法和可操作的控制措施,提升AI系统的可信度和合规性。

Comments We have adjustments to make for higher effectiveness

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07806 2026-01-13 cs.CL cs.LG 62%

The Confidence Trap: Gender Bias and Predictive Certainty in LLMs

自信陷阱:大语言模型中的性别偏见与预测确定性

Ahmed Sabir, Markus Kängsepp, Rajesh Sharma

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 本研究探讨大语言模型在性别偏见任务中的置信度校准问题,提出Gender-ECE指标以评估公平性差异,发现Gemma-2在性别偏见基准中校准最差。

Comments AAAI 2026 (AISI Track), Oral. Project page: https://bit.ly/4p8OKQD

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06223 2026-01-13 cs.CY cs.AI 62%

Toward Safe and Responsible AI Agents: A Three-Pillar Model for Transparency, Accountability, and Trustworthiness

迈向安全和负责任的AI代理:透明、问责和可信的三支柱模型

Edward C. Cheng, Jeshua Cheng, Alice Siu

机构 * InquiryOn

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

AI总结 本文提出三支柱模型,旨在通过透明、问责和可信机制,确保AI代理的安全性和责任性,促进负责任的AI发展。

Comments 15 pages, 8 figures, conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06101 2026-01-13 cs.CY cs.AI 62%

How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures

如何评估人工智能素养:自我报告与基于客观的评估之间的不一致

Shan Zhang, Ruiwei Xiao, Anthony F. Botelho, Guanze Liao, Thomas K. F. Chiu, John Stamper, Kenneth R. Koedinger

机构 * University of Florida(佛罗里达大学) Carnegie Mellon University(卡内基梅隆大学) National Tsing Hua University(国立清华大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本研究开发并评估了教师人工智能素养的自我报告和基于客观的测量方法,揭示了两者在不同教师群体中的显著差异,并提出了适用于专业发展的诊断工具。

Comments 16 pages, 6 figures, LAK2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06064 2026-01-13 cs.CY cs.AI cs.MA 62%

Socio-technical aspects of Agentic AI

群体技术视角下的代理AI

Praveen Kumar Donta, Alaa Saleh, Ying Li, Shubham Vaishnav, Kai Fang, Hailin Feng, Yuchao Xia, Thippa Reddy Gadekallu, Qiyang Zhang, Xiaodan Shi, Ali Beikmohammadi, Sindri Magnússon, Ilir Murturi, Chinmaya Kumar Dehury, Marcin Paprzycki, Lauri Loven, Sasu Tarkoma, Schahram Dustdar

机构 * Department of Computer and Systems Sciences, Stockholm University(斯德哥尔摩大学计算机与系统科学系) Center for Ubiquitous Computing, University of Oulu(奥卢大学无处不在计算中心) College of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院) Zhejiang A\&F University, Hangzhou(浙江工业大学之江学院) School of Computer Science, Peking University(北京大学计算机科学学院) Department of Mechatronics, University of Prishtina(普里什蒂纳大学机电系) Department of Computer Science, IISER Berhampur(伯尔哈普尔IISER计算机科学系) Systems Research Institute Polish Academy of Sciences(波兰科学院系统研究所) Department of Computer Science, University of Helsinki(赫尔辛基大学计算机科学系)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 本文从社会技术视角探讨代理AI,分析其技术组件与社会背景的关联,揭示伦理挑战及未来研究方向。

Comments Dear Reviewer, please note that this is not survey/review or position paper. This paper introduced new framework (MAD-BAD-SAD Framework) for Socio-technical aspects of Agentic AI, Ethical considerations, which is very important to consider beside technical development

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05641 2026-01-12 cs.CL cs.LG 62%

Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs

多语言失忆:多语言大语言模型中卸载的可转移性研究

Alireza Dehghanpour Farashah, Aditi Khandelwal, Marylou Fauchard, Zhuan Shi, Negar Rostamzadeh, Golnoosh Farnadi

机构 * Mila – Quebec AI Institute(魁北克人工智能研究所) McGill University(麦吉尔大学) Université de Montréal(蒙特利尔大学) Google Research(谷歌研究)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.LG

AI总结 本研究探讨多语言大语言模型中卸载的可转移性,通过实验发现语法相似性是预测跨语言卸载行为的关键因素。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05184 2026-01-09 cs.AI cs.CL 62%

Observations and Remedies for Large Language Model Bias in Self-Consuming Performative Loop

大语言模型自我消耗表现性循环中的观测与对策

Yaxuan Wang, Zhongteng Cai, Yujia Bao, Xueru Zhang, Yang Liu

机构 * University of California, Santa Cruz(加州大学圣克ruz分校) The Ohio State University(俄亥俄州立大学) Center for Advanced AI, Accenture(Accenture高级人工智能中心)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出自我消耗表现性循环(SCPL)概念,通过实验发现表现性循环会增加偏好偏见并降低差异偏见,设计奖励拒绝采样策略以缓解偏见。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04897 2026-01-09 cs.CL cs.CV cs.LG cs.MM 62%

V-FAT: Benchmarking Visual Fidelity Against Text-bias

V-FAT:基于文本偏见的视觉真实性基准测试

Ziteng Wang, Yujie He, Guanliang Li, Siqi Yang, Jiaqi Xiong, Songxiang Liu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Meituan(美团) University of Oxford(牛津大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG

AI总结 V-FAT通过三级评估框架量化文本偏见对视觉真实性的干扰,揭示多模态大语言模型在高语言主导性下的视觉崩溃问题。

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04747 2026-01-08 cs.CY cs.AI 62%

E-LENS: User Requirements-Oriented AI Ethics Assurance

E-LENS:以用户需求为导向的AI伦理保证

Jianlong Zhou, Fang Chen

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 本文提出以用户需求为导向的AI伦理保证方法,通过E-LENS平台验证AI伦理,以航空安全方法为借鉴,结合危险分析构建AI伦理保证案例。

Comments 29 pages

Journal ref Human-Intelligent Systems Integration, August 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00927 2026-01-06 cs.SI cs.AI cs.CL 62%

Measuring Social Media Polarization Using Large Language Models and Heuristic Rules

利用大语言模型和启发规则测量社交媒体极化

Jawad Chowdhury, Rezaur Rashid, Gabriel Terejanu

机构 * Department of Computer Science, University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校计算机科学系)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文提出利用大语言模型和启发规则分析社交媒体讨论中的情感极化现象,揭示事件驱动的极化模式并提供可扩展的量化方法。

Comments Foundations and Applications of Big Data Analytics (FAB), Niagara Falls, Canada, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00816 2026-01-06 cs.AI cs.CR cs.LG 62%

MathLedger: A Verifiable Learning Substrate with Ledger-Attested Feedback

MathLedger: 一种具有账本证明反馈的可验证学习基础

Ismail Ahmad Abdullah

机构 * CNU(中国矿业大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

AI总结 MathLedger通过整合形式验证、密码学证明和学习动态,提供一种可验证学习的基础,实现可审计的机器认知系统。

Comments 14 pages, 1 figure, 2 tables, 2 appendices with full proofs. Documents v0.9.4-pilot-audit-hardened audit surface with fail-closed governance, canonical JSON hashing, and artifact classification. Phase I infrastructure validation; no capability claims

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22725 2025-12-30 cs.CL cs.CY 62%

Mitigating Social Desirability Bias in Random Silicon Sampling

在随机硅采样中缓解社会可取性偏差

Sashank Chapala, Maksym Mironov, Songgaojun Deng

机构 * Eindhoven University of Technology(埃因霍温理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.CY

AI总结 本研究通过提示工程方法减轻LLM中的社会可取性偏差,提升硅样本与人类数据的一致性。

Comments 31 pages, 9 figures, and 24 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21775 2025-12-29 cs.AI cs.CY cs.DB 62%

Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets

合规评分方案:一种面向生成式人工智能数据集的数据溯源框架

Matyas Bohacek, Ignacio Vilanova Echavarri

机构 * Stanford University(斯坦福大学) Imperial College London(帝国理工学院伦敦分校)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 本文提出合规评分方案,用于评估生成式人工智能数据集的透明度、问责制和安全性,同时提供开源库以实现数据溯源,促进负责任的数据集构建与管理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13540 2025-12-29 cs.LG cs.CY 62%

Fairness-Aware Graph Representation Learning with Limited Demographic Information

带有有限人口信息的公平图表示学习

Zichong Wang, Zhipeng Yin, Liping Yang, Jun Zhuang, Rui Yu, Qingzhao Kong, Wenbin Zhang

机构 * Florida International University(佛罗里达国际大学) University of New Mexico(新墨西哥大学) Boise State University(博伊西州立大学) University of Louisville(路易斯维尔大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.CY、cs.LG

AI总结 本文提出FairGLite框架,在有限人口信息下减轻图学习中的偏见,通过生成人口信息代理和自适应置信度策略,实现公平性和效用的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00931 2025-12-19 cs.CL cs.AI 62%

Mitigating Hallucinations in Zero-Shot Scientific Summarisation: A Pilot Study

缓解零样本科学摘要中的幻觉:一项初步研究

Imane Jaaouine, Ross D. King

机构 * Department of Chemical Engineering and Biotechnology(化学工程与生物技术系) University of Cambridge(剑桥大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本研究探讨提示工程方法对缓解零样本科学摘要中的幻觉效果,发现上下文重复和随机添加能显著提升摘要与原文的对齐度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13702 2025-12-17 cs.CY cs.AI 62%

Enhancing Transparency and Traceability in Healthcare AI: The AI Product Passport

提升医疗AI的透明度与可追溯性:AI产品护照

A. Anil Sinaci, Senan Postaci, Dogukan Cavdaroglu, Machteld J. Boonstra, Okan Mercan, Kerem Yilmaz, Gokce B. Laleci Erturkmen, Folkert W. Asselbergs, Karim Lekadir

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本研究提出AI产品护照,通过生命周期文档提升医疗AI的透明度和可追溯性,符合FUTURE-AI原则,确保公平性和可用性,并通过开源平台实现可访问性。

Comments A total of 33 pages: First 16 pages for the manuscript and the remaining 17 pages for the supplementary user guide of the graphical user interface

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16069 2025-12-12 cs.CY cs.AI 62%

Human or AI? Comparing Design Thinking Assessments by Teaching Assistants and Bots

人类还是人工智能?教学助教与机器人的设计思维评估比较

Sumbul Khan, Wei Ting Liow, Lay Kee Ang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本研究比较了教学助教与AI在评估设计思维教育学生海报中的表现,发现教师更偏好助教评分,但AI在一致性与反馈效率上具有一定优势。

Comments to be published in IEEE TALE 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09483 2025-12-11 cs.CL cs.CY 62%

Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines

基于大语言模型的搜索引擎与传统搜索引擎的来源覆盖与引用偏见

Peixian Zhang, Qiming Ye, Zifan Peng, Kiran Garimella, Gareth Tyson

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Rutgers University(罗格斯大学) Rutgers University New Brunswick United States(罗格斯大学新 Brunswick美国)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.CY

AI总结 本文研究了基于大语言模型的搜索引擎与传统搜索引擎在来源覆盖和引用偏见方面的差异,发现LLM-SEs在资源多样性上优于传统搜索引擎,但其可信度和中立性仍需进一步提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09458 2025-12-11 cs.AI cs.LG 62%

Architectures for Building Agentic AI

构建代理AI的架构

Sławomir Nowaczyk

机构 * Center for Applied Intelligent Systems Research, Halmstad University, Sweden(应用智能系统研究所,哈尔姆斯塔德大学,瑞典)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种构建代理AI的架构分类,通过组件化设计和显式控制循环提升系统可靠性。

Comments This is a preprint of a chapter accepted for publication in Generative and Agentic AI Reliability: Architectures, Challenges, and Trust for Autonomous Systems, published by Springer Nature

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08978 2025-12-11 cs.CY cs.AI 62%

Institutional AI Sovereignty Through Gateway Architecture: Implementation Report from Fontys ICT

通过网关架构实现机构AI主权:Fontys ICT的实施报告

Ruud Huijts, Koen Suilen

机构 * Fontys ICT

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出通过网关架构实现机构AI主权,构建了受监管的AI平台,实现可控的AI访问和治理,强调AI作为战略工具需专门领导和治理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08592 2025-12-10 cs.AI cs.CY cs.HC cs.SY eess.SY 62%

The SMART+ Framework for AI Systems

为AI系统设计的SMART+框架

Laxmiraju Kandikatla, Branislav Radeljic

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 SMART+框架为AI系统提供了一种结构化模型,涵盖安全、监控、责任、可靠性和透明性,并增强隐私与安全、数据治理和公平性,以提升AI系统的治理和合规性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04691 2025-12-05 cs.AI cs.CL cs.MA 62%

Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective

迈向伦理化的大型语言模型多智能体系统:从机制可解释性视角出发

Jae Hee Lee, Anne Lauscher, Stefano V. Albrecht

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文从机制可解释性视角出发,提出确保大型语言模型多智能体系统伦理行为的研究议程,聚焦于伦理评估框架、内部机制解析及参数高效对齐技术。

Comments Accepted to LaMAS 2026@AAAI'26 (https://sites.google.com/view/lamas2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23733 2025-12-04 cs.CY cs.AI cs.HC 62%

Unintentional Consequences: Generative AI Use for Cybercrime

无意后果:生成式AI用于网络犯罪

Truong Jack Luu, Binny M. Samuel

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 生成式AI的普及导致网络犯罪激增,研究通过分析数据揭示AI技术放大恶意行为的机制,并提出多层策略以应对风险。

详情

展开后加载摘要…

URL PDF HTML 收藏