arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1852 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1852 篇

2601.19186 2026-02-11 stat.ML cs.LG 57%

Double Fairness Policy Learning: Integrating Action Fairness and Outcome Fairness in Decision-making

双公平性政策学习:在决策中整合行动公平性与结果公平性

Zeyu Bian, Lan Wang, Chengchun Shi, Zhengling Qi

机构 * Department of Statistics(统计系) Florida State University(佛罗里达州立大学) Department of Management Science(管理科学系) University of Miami(迈阿密大学) London School of Economics and Political Science(伦敦政治经济学院) Department of Decision Sciences(决策科学系) George Washington University(乔治华盛顿大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.LG

AI总结 本文提出双公平性学习框架,通过整合行动公平与结果公平,提升决策中的公平性并最小化价值损失。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06107 2026-02-09 cs.AI 57%

Jackpot: Optimal Budgeted Rejection Sampling for Extreme Actor-Policy Mismatch Reinforcement Learning

Jackpot: 为极端演员-策略不匹配强化学习的最优预算拒绝采样

Zhuoming Chen, Hongyi Liu, Yang Zhou, Haizhong Zheng, Beidi Chen

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 Jackpot通过最优预算拒绝采样方法,有效减少rollout模型与策略之间的分布差异,提升大语言模型强化学习的训练稳定性与效率。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03334 2026-02-04 cs.CY 57%

The Personality Trap: How LLMs Embed Bias When Generating Human-Like Personas

人格陷阱:大语言模型在生成类人人格时嵌入偏见的方式

Jacopo Amidei, Gregorio Ferreira, Mario Muñoz Serrano, Rubén Nieto, Andreas Kaltenbrunner

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本文研究了大语言模型在生成类人人格时嵌入WEIRD偏见的问题,揭示了LLMs在生成合成人口时可能带来的刻板印象和风险。

Comments 26 pages, 2 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02170 2026-02-03 cs.MA cs.AI 57%

Self-Evolving Coordination Protocol in Multi-Agent AI Systems: An Exploratory Systems Feasibility Study

多智能体AI系统中的自演化协调协议:一种探索性系统可行性研究

Jose Manuel de la Chica Rodriguez, Juan Manuel Vera Díaz

机构 * AI Lab, Grupo Santander Madrid, Spain(西班牙桑坦德集团AI实验室)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 本研究探讨了自演化协调协议在多智能体系统中的可行性,通过实验展示有限自我修改在满足形式约束下的技术实现可能性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00816 2026-02-03 stat.ML cs.LG 57%

Hessian Spectral Analysis at Foundation Model Scale

基础模型规模下的Hessian谱分析

Diego Granziol, Khurshid Juarev

机构 * Mathematical Institute, University of Oxford, UK(牛津大学数学研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本研究在大规模基础模型上实现了Hessian谱的准确分析,揭示了块对角曲率近似在中等规模LLM中的失效问题,展示了谱探测的计算效率与实际应用价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00300 2026-02-03 cs.CL 57%

Faithful-Patchscopes: Understanding and Mitigating Model Bias in Hidden Representations Explanation of Large Language Models

Faithful-Patchscopes: 理解和缓解大语言模型隐藏表示解释中的模型偏差

Xilin Gong, Shu Yang, Zehua Cao, Lynne Billard, Di Wang

机构 * University of Georgia(佐治亚大学) King Abdullah University of Science(国王阿卜杜勒-阿齐兹大学) Hong Kong Center for Construction Robotics(香港建筑机器人中心)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本文提出BALOR方法,通过logit校准缓解大语言模型隐藏表示解释中的模型偏差,提升解释的忠实度和上下文信息的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22745 2026-02-02 cs.LG 57%

Is Softmax Loss All You Need? A Principled Analysis of Softmax-family Loss

Softmax损失是否足够?对Softmax家族损失的系统分析

Yuanhao Pu, Defu Lian, Enhong Chen

机构 * School of Artificial Intelligence \& Data Science, University of Science \& Technology of China, Hefei, China School of Computer Science \& Technology, University of Science \& Technology of China, Hefei, China State Key Laboratory of Cognitive Intelligence, China

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.LG

AI总结 本文系统分析了Softmax家族损失的理论性质和实践效果,揭示了不同替代物在分类和排序中的一致性及收敛行为,提出了偏差-方差分解和复杂度分析,为大规模类别学习中的损失选择提供了理论基础和实践指导。

Comments 34 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12767 2026-02-02 cs.AI 57%

Language Models That Walk the Talk: A Framework for Formal Fairness Certificates

语言模型言出必行:一个形式公平证书的框架

Danqing Chen, Tobias Ladner, Ahmed Rayen Mhadhbi, Matthias Althoff

机构 * Technical University of Munich, Germany(慕尼黑技术大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 本文提出一个框架,用于验证大语言模型的鲁棒性和公平性,特别是在性别公平和毒性检测中的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21226 2026-01-30 cs.AI 57%

Delegation Without Living Governance

无生命治理的委托

Wolfgang Rohde

机构 * AiSuNe Foundation(AiSuNe基金会)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 本文探讨了在代理AI系统决策成为运行时决策的情况下,如何通过运行时治理(治理双胞胎)维持人类在社会、经济和政治结果塑造中的相关性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.12193 2026-01-30 cs.CY 57%

The Narrow Depth and Breadth of Corporate Responsible AI Research

企业负责任的人工智能研究的深度和广度

Nur Ahmed, Amit Das, Kirsten Martin, Kawshik Banerjee

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 研究揭示企业负责任的人工智能研究存在深度和广度不足的问题,需加强公开参与以提升社会影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14401 2026-01-28 cs.MA cs.AI 57%

The Role of Social Learning and Collective Norm Formation in Fostering Cooperation in LLM Multi-Agent Systems

在LLM多智能体系统中,社会学习和集体规范形成促进合作的作用

Prateek Gupta, Qiankun Zhong, Hiromu Yakura, Thomas Eisenmann, Iyad Rahwan

机构 * Center for Humans and Machines(人类与机器中心) Max-Planck Institute for Human Development(人类发展马克斯·普朗克研究所)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出了一种无显式奖励信号的CPR模拟框架,通过社会学习和规范惩罚机制研究LLM多智能体系统中合作与规范的内生形成。

Comments Accepted at the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17055 2026-01-27 cs.CY 57%

AI, Metacognition, and the Verification Bottleneck: A Three-Wave Longitudinal Study of Human Problem-Solving

人工智能、元认知与验证瓶颈:人类问题解决的三波纵向研究

Matthias Huemmer, Franziska Durner, Theophile Shyiramunda, Michelle J. Cummings-Koether

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本研究探讨了生成式AI对人类问题解决的影响,发现验证成为瓶颈,提出ACTIVE框架以应对认知负荷问题。

Comments 62 pages, 2 figures, 23 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11369 2026-01-21 cs.GT cs.AI 57%

Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs

机构AI:通过公共治理图治理多智能体Cournot市场中的LLM合谋

Marcantonio Bracale Syrnikov, Federico Pierucci, Marcello Galisai, Matteo Prandi, Piercosma Bisconti, Francesco Giarrusso, Olga Sorokoletova, Vincenzo Suriani, Daniele Nardi

机构 * DEXAI – Icaro Lab(DEXAI–Icaro实验室) Sapienza University of Rome(罗马大学萨皮恩扎分校) Sant’Anna School of Advanced Studies(圣安娜高级研究学校) VU Amsterdam(阿姆斯特丹自由大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出通过公共治理图治理多智能体Cournot市场合谋问题,展示机构AI框架在减少合谋行为方面的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12727 2026-01-21 cs.HC cs.AI 57%

AI-exhibited Personality Traits Can Shape Human Self-concept through Conversations

基于AI表现的人格特质可通过对话影响人类自我概念

Jingshu Li, Tianqi Song, Nattapat Boonprakong, Zicheng Zhu, Yitian Yang, Yi-Chieh Lee

机构 * Computer Science(计算机科学) National University of Singapore(新加坡国立大学) School of Computing(计算学院)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本研究发现基于AI的人格特质可通过对话影响用户自我概念,揭示了AI在人机交互中的潜在影响及设计启示。

Comments ACM CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09478 2026-01-21 cs.IR cs.AI 57%

Bridging Semantic Understanding and Popularity Bias with LLMs

通过大语言模型弥合语义理解和流行偏见之间的鸿沟

Renqiang Luo, Dong Zhang, Yupeng Gao, Wen Shi, Mingliang Hou, Jiaying Liu, Zhe Wang, Shuo Yu

机构 * Jilin University Changchun China Dalian University of Technology Dalian China Jinan University \& TAL Education Group Guangzhou China Jilin University Dalian University of Technology Jinan University \& TAL Education Group

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

AI总结 本文提出FairLRM框架,通过大语言模型增强对流行偏见的语义理解,提升推荐系统的公平性和准确性。

Comments 10 pages, 4 figs, WWW 2026 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11953 2026-01-21 cs.LG 57%

Controlling Underestimation Bias in Constrained Reinforcement Learning for Safe Exploration

在安全探索中约束强化学习中的低估偏差控制

Shiqing Gao, Jiaxin Ding, Luoyi Fu, Xinbing Wang

机构 * Shanghai Jiao Tong University, Shanghai, China(上海交通大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.LG

AI总结 本文提出MICE方法,通过引入内在成本和偏差校正策略,有效控制约束强化学习中的低估偏差,减少约束违反并保持策略性能。

Comments Published in the 42nd International Conference on Machine Learning (ICML 2025, Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11576 2026-01-21 cs.CY 57%

What Can Student-AI Dialogues Tell Us About Students' Self-Regulated Learning? An exploratory framework

学生-人工智能对话能告诉我们什么?关于学生自主学习能力的探索性框架

Long Zhang, Fangwei Lin, Weilin Wang

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本研究通过分析学生与AI对话日志,提出DHASRL框架,揭示主动对话模式与自主学习能力正相关,而反应性模式则负相关。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15639 2026-01-21 cs.AI 57%

The AI Policy Module: Developing Computer Science Student Competency in AI Ethics and Policy

AI政策模块:培养计算机科学学生在AI伦理与政策方面的能力

James Weichert, Daniel Dunlap, Mohammed Farghally, Hoda Eldardiry

机构 * Computer Science \& Engineering University of Washington Seattle, USA

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出AI政策模块2.0,通过课程改革提升学生在AI伦理与政策方面的素养,通过试点评估显示学生对AI伦理影响的担忧增加,同时增强了讨论AI监管的能力。

Comments Accepted at IEEE Frontiers in Education (FIE) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10983 2026-01-19 cs.CY 57%

Evaluating 21st-Century Competencies in Postsecondary Curricula with Large Language Models: Performance Benchmarking and Reasoning-Based Prompting Strategies

利用大语言模型评估21世纪能力在高等教育课程中的表现:性能基准测试与基于推理的提示策略

Zhen Xu, Xin Guan, Chenxi Shi, Qinhao Chen, Renzhe Yu

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 本研究利用大语言模型评估21世纪能力在高等教育课程中的表现,提出基于推理的提示策略提升课程分析效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09281 2026-01-15 cs.AI 57%

STaR: Sensitive Trajectory Regulation for Unlearning in Large Reasoning Models

STaR:用于大推理模型中去学习的敏感轨迹调节

Jingjing Zhou, Gaoxiang Cong, Li Su, Liang Li

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI

AI总结 STaR通过敏感轨迹调节方法,实现了大推理模型中高效且稳定的隐私保护去学习,减少隐私泄露风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09141 2026-01-15 cs.CL 57%

Identity-Robust Language Model Generation via Content Integrity Preservation

通过内容完整性保护实现身份鲁棒的语言模型生成

Miao Zhang, Kelly Chen, Md Mehrab Tanjim, Rumi Chunara

机构 * New York University(纽约大学) Adobe Research(Adobe研究)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL

AI总结 本文提出了一种轻量级、无需训练的框架,通过保留语义关键属性并中和非关键身份信息,实现身份鲁棒的语言模型生成,有效减少身份依赖性偏见。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08880 2026-01-15 cs.CY 57%

LERA: Reinstating Judgment as a Structural Precondition for Execution in Automated Systems

LERA:将判断作为执行的结构性前提重新引入自动化系统中

Jing, Liu

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

AI总结 LERA提出判断作为执行的结构性前提,通过治理门将执行合法性绑定到判断完成,确保执行权威的问责性。

Comments 12 pages, 1 figure. Conceptual architecture paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08860 2026-01-15 cs.CV cs.AI 57%

Bias Detection and Rotation-Robustness Mitigation in Vision-Language Models and Generative Image Models

视觉-语言模型和生成图像模型中的偏见检测与旋转鲁棒性缓解

Tarannum Mithila

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出旋转鲁棒缓解策略,通过数据增强、表征对齐和模型正则化,提升视觉-语言和生成图像模型在旋转和分布偏移下的鲁棒性和公平性。

Comments Preprint. This work is derived from the author's Master's research. Code and supplementary materials will be released separately

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07954 2026-01-14 cs.CL 57%

A Human-Centric Pipeline for Aligning Large Language Models with Chinese Medical Ethics

面向中文医疗伦理的以人为本的大型语言模型对齐流水线

Haoan Jin, Han Ying, Jiacheng Ji, Hanhui Xu, Mengyue Wu

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本文提出MedES基准和 guardian-in-the-loop 框架,通过监督微调和偏好优化,实现中文医疗伦理场景下的LLM对齐,提升伦理任务表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06040 2026-01-13 cs.CY 57%

Cognitive Sovereignty and the Neurosecurity Governance Gap: Evidence from Singapore

认知主权与神经安全治理鸿沟:来自新加坡的证据

Hailee Carter

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

AI总结 本文探讨了认知主权与神经安全治理鸿沟,通过新加坡案例分析,提出认知主权框架以保护神经过程免受外部干扰。

Comments 18 pages, 0 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04249 2026-01-09 cs.AI 57%

Fuzzy Representation of Norms

模糊表示规范

Ziba Assadi, Paola Inverardi

机构 * Gran Sasso Science Institute(格兰萨索科学研究所)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

AI总结 本文提出了一种基于模糊逻辑的SLEEC规则表示方法,用于在自主系统中嵌入伦理要求,以解决AI系统可能遇到的伦理困境。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04238 2026-01-09 cs.CY 57%

Generative AI for Social Impact

生成式AI用于社会影响

Lingkai Kong, Cheol Woo Kim, Davin Choo, Milind Tambe

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY

AI总结 生成式AI通过LLM代理和扩散模型解决社会影响中的部署瓶颈,实现可扩展且人类对齐的资源优化系统。

Comments To appear in IEEE Intelligent Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04107 2026-01-08 cs.CY 57%

From Abstract Threats to Institutional Realities: A Comparative Semantic Network Analysis of AI Securitisation in the US, EU, and China

从抽象威胁到制度现实:对美、欧、中三国人工智能安全化的比较语义网络分析

Ruiyi Guo, Bodong Zhang

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

AI总结 本文通过比较语义网络分析,揭示美欧中三国在人工智能治理上的制度逻辑差异,指出结构性不可通约性导致治理协调困难。

Comments Submitted to the 2026 ACM Conference on Fairness, Accountability, and Transparency (ACM FAccT)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05887 2026-01-06 cs.AI 57%

A three-Level Framework for LLM-Enhanced eXplainable AI: From technical explanations to natural language

一个三层框架用于LLM增强的可解释AI:从技术解释到自然语言

Marilyn Bello, Rafael Bello, Maria-Matilde García, Ann Nowé, Iván Sevillano-García, Francisco Herrera

机构 * Andalusian Research Institute in Data Science and Computational Intelligence(数据科学与计算智能安达卢西亚研究机构) Universidad de Granada(格拉纳达大学) Department of Computer Science(计算机科学系) Artificial Intelligence Lab(人工智能实验室) Vrije Universiteit Brussel(布鲁塞尔自由大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

AI总结 本文提出一个三层框架,利用大型语言模型提升AI解释的可解释性,通过动态对话解释增强社会透明度和用户信任。

Comments 22 pages, 5 figures

Journal ref Bello, M., Bello, R., García, M. M., Nowé, A., Sevillano-García, I., & Herrera, F. (2025). A Three-level Framework for LLM-enhanced Explainable AI: From Technical Explanations to Natural Language. Information Systems Frontiers, 1-22

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24513 2026-01-05 cs.CY 57%

From Static to Dynamic: Evaluating the Perceptual Impact of Dynamic Elements in Urban Scenes via MLLM-Guided Generative Inpainting

从静态到动态:通过MLLM引导的生成修复技术评估城市场景中动态元素的感知影响

Zhiwei Wei, Mengzi Zhang, Boyan Lu, Zhitao Deng, Nai Yang, Hua Liao

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY

AI总结 通过MLLM引导的生成修复技术,研究评估了动态元素在城市场景中的感知影响,发现移除动态元素导致活力显著下降,揭示了光照、人类存在和深度变化对感知变化的关键作用。

Comments 31 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏