arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1847 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1847 篇

2209.13020 2023-05-17 cs.CY cs.AI cs.LG 67%

Law Informs Code: A Legal Informatics Approach to Aligning Artificial Intelligence with Humans

John J. Nay

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Northwestern Journal of Technology and Intellectual Property, Volume 20, Issue 3, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.15010 2023-05-01 cs.CV cs.AI cs.CL cs.LG cs.MM 67%

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, Hongsheng Li, Yu Qiao

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Code and models are available at https://github.com/ZrrSkywalker/LLaMA-Adapter

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.06346 2022-11-14 cs.CY cs.AI cs.LG 67%

AI Ethics in Smart Healthcare

Sudeep Pasricha

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.00504 2022-10-18 stat.ML cs.AI cs.CY cs.LG 67%

Domain Adaptation meets Individual Fairness. And they get along

Debarghya Mukherjee, Felix Petersen, Mikhail Yurochkin, Yuekai Sun

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Published at NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.11569 2022-07-26 cs.RO cs.AI cs.CV cs.CY cs.LG 67%

Robots Enact Malignant Stereotypes

Andrew Hundt, William Agnew, Vicky Zeng, Severin Kacianka, Matthew Gombolay

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 30 pages, 10 figures, 5 tables. Website: https://sites.google.com/view/robots-enact-stereotypes . Published in the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT 22), June 21-24, 2022, Seoul, Republic of Korea. ACM, DOI: https://doi.org/10.1145/3531146.3533138 . FAccT22 Submission dates: Abstract Dec 13, 2021; Submitted Jan 22, 2022; Accepted Apr 7, 2022

Journal ref In 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT 22). ACM, New York, NY, USA, 743-756

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.03295 2022-06-03 cs.LG cs.AI cs.CY 67%

The Road to Explainability is Paved with Bias: Measuring the Fairness of Explanations

Aparna Balagopalan, Haoran Zhang, Kimia Hamidieh, Thomas Hartvigsen, Frank Rudzicz, Marzyeh Ghassemi

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Published in FAccT 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.09447 2022-06-02 cs.LG cs.AI cs.CY stat.AP 67%

Algorithmic Fairness Verification with Graphical Models

Bishwamittra Ghosh, Debabrota Basu, Kuldeep S. Meel

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.05151 2022-04-12 cs.CY cs.AI cs.LG 67%

Metaethical Perspectives on 'Benchmarking' AI Ethics

Travis LaCroix, Alexandra Sasha Luccioni

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments 39 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.10476 2022-03-01 cs.LG cs.AI cs.CY stat.ML 67%

Towards Return Parity in Markov Decision Processes

Jianfeng Chi, Jian Shen, Xinyi Dai, Weinan Zhang, Yuan Tian, Han Zhao

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments AISTATS 2022. Code is released at https://github.com/JFChi/Return-Parity-MDP

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.12305 2021-07-23 cs.CL cs.AI cs.CY 67%

Confronting Abusive Language Online: A Survey from the Ethical and Human Rights Perspective

Svetlana Kiritchenko, Isar Nejadgholi, Kathleen C. Fraser

专题命中 AI治理与伦理 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments published in Journal of Artificial Intelligence Research, 71: 431-478, July 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.04257 2021-04-29 cs.CY cs.AI cs.LG 67%

Fairness for Unobserved Characteristics: Insights from Technological Impacts on Queer Communities

Nenad Tomasev, Kevin R. McKee, Jackie Kay, Shakir Mohamed

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society (AIES 2021)

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.01468 2020-06-09 cs.LG cs.AI cs.CY stat.ML 67%

Auditing and Achieving Intersectional Fairness in Classification Problems

Giulio Morina, Viktoriia Oliinyk, Julian Waton, Ines Marusic, Konstantinos Georgatzis

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11551 2025-12-09 cs.AI cs.CL 66%

Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping

对 Machiavellian 代理进行对齐:通过测试时策略塑造实现行为引导

Dena Mujtaba, Brian Hu, Anthony Hoogs, Arslan Basharat

专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CL、cs.AI

AI总结 本文提出了一种测试时策略塑造方法,通过模型引导的策略调整,解决预训练代理在复杂环境中的伦理对齐问题,实现奖励最大化与伦理约束的平衡。

Comments Accepted to AAAI 2026 AI Alignment Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10590 2025-11-13 cs.CY cs.AI 66%

Machine Unlearning for Responsible and Adaptive AI in Education

Betty Mayeku, Sandra Hummel, Parisa Memarmoshrefi

机构 * Leipzig University(莱比锡大学) Technical University Dresden(德累斯顿技术大学) University of Göttingen(哥廷根大学)

专题命中 AI治理与伦理 :trustworthy(abstract,comments);分类 cs.AI、cs.CY

Comments Accepted paper - ESORICS 2025 - International Workshop on Secure and Trustworthy Machine Unlearning Systems (STMUS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18562 2025-10-07 cs.CL cs.AI 66%

From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test

Xunlian Dai, Li Zhou, Benyou Wang, Haizhou Li

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院)

专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.CL、cs.AI

Comments Cultural Analysis, Cultural Alignment, Word Association Test, Large Language Models. Accepted by EMNLP 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10073 2025-08-01 cs.CL cs.AI 66%

Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires

Simon Münker

机构 * Tier University(Tier大学)

专题命中 AI治理与伦理 :alignment(abstract,journal_ref);分类 cs.CL、cs.AI

Comments 15pages, 1 figure, 2 tables

Journal ref Proceedings of 0th Symposium on Moral and Legal AI Alignment of the IACAP/AISB Conference, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13533 2025-01-24 cs.AI cs.LG 66%

Towards a Theory of AI Personhood

Francis Rhys Ward

专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.AI、cs.LG

Comments AAAI-25 AI Alignment Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.16872 2024-12-09 cs.AI cs.LG 66%

Ethical and Scalable Automation: A Governance and Compliance Framework for Business Applications

Haocheng Lin

专题命中 AI治理与伦理 :alignment(abstract,comments);分类 cs.AI、cs.LG

Comments The current version improves significantly by integrating ethical frameworks, expanding methodology and case studies, enhancing scalability and ethical-legal alignment, acknowledging prior work, and offering clearer structure and practical relevance

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01760 2024-08-28 cs.CY cs.AI 66%

Trust and ethical considerations in a multi-modal, explainable AI-driven chatbot tutoring system: The case of collaboratively solving Rubik's Cube

Kausik Lakkaraju, Vedant Khandelwal, Biplav Srivastava, Forest Agostinelli, Hengtao Tang, Prathamjeet Singh, Dezhi Wu, Matt Irvin, Ashish Kundu

专题命中 AI治理与伦理 :trustworthy(abstract,comments);分类 cs.AI、cs.CY

Comments Accepted at 'Neural Conversational AI Workshop - What's left to TEACH (Trustworthy, Enhanced, Adaptable, Capable, and Human-centric) chatbots?' at ICML 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18281 2026-05-19 cs.LG 65%

Temporal Task Diversity: Inductive Biases Under Non-Stationarity in Synthetic Sequence Modelling

时间任务多样性:非平稳性下的归纳偏置

Afiq Abdillah Effiezal Aswadi, Oliver Britton, Ross Baker, Matthew Farrugia-Roberts

机构 * University of Oxford(牛津大学)

专题命中 AI治理与伦理 :safety(abstract,comments);分类 cs.LG;AI safety(comments)

AI总结 研究探讨了在合成序列建模中,任务分布随时间变化对深度学习模型归纳偏置的影响,发现任务分布的多样性增强了模型对泛化而非记忆的偏好。

Comments Presented at Technical AI Safety Conference (TAIS), Oxford, May 2026. Code available at https://github.com/matomatical/temporal-task-diversity

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.21577 2026-08-25 cs.LG cs.AI 新提交 62%

Anchoring Bias: A Persistent Fairness Backdoor Attack against MLLMs under Continual Learning

锚定偏差:一种针对持续学习下多模态大语言模型(MLLMs)的持久公平性后门攻击

Yuyang Luo, Kai Shu

机构 * Emory University(埃默里大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

AI总结 针对持续学习下的多模态大语言模型,研究人员提出持久公平性后门攻击,通过两种机制注入持久群体歧视,该攻击能规避标准防御且在多轮持续学习中留存。

Comments CIKM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10169 2026-08-25 cs.AI cs.LG 版本更新 62%

MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction

MAVEN-T:用于实时多智能体轨迹预测的强化异构蒸馏

Wenchang Duan, Zhenguo Gao, Jinguo Xian, Yi Shi

机构 * School of Mathematical Sciences, Shanghai Jiao Tong University(上海交通大学数学科学学院) Bio-X Institutes, Key Laboratory for the Genetics of Developmental and Neuropsychiatric Disorders, Shanghai Jiao Tong University(上海交通大学Bio-X研究院、发育与神经精神疾病遗传学重点实验室) Shanghai Key Laboratory of Psychotic Disorders, Brain Science and Technology Research Center, Shanghai Jiao Tong University(上海精神疾病重点实验室、脑科学与技术研究中心,上海交通大学)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

AI总结 提出MAVEN-T框架,通过高容量教师模型和紧凑学生模型的异构蒸馏,结合强化学习优化,实现实时多智能体轨迹预测,在多个数据集上达到高精度与低延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20421 2026-08-24 cs.CY cs.AI 新提交 62%

Six misconceptions about large language models: A minimal model and diagnostic taxonomy

大型语言模型的六大误解:一个极简模型与诊断分类法

Zhicheng Lin

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 该研究提出以四组区分为核心的LLM极简模型,诊断六大误解,应用于出版商AI政策案例,为纠正民间理论错误提供诊断工具。

Comments 20 pages, 1 figure, 2 tables, and 2 boxes. Published in PNAS Nexus

Journal ref PNAS Nexus, 5(7), pgag236 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18122 2026-08-20 cs.CY cs.AI 新提交 62%

Global Index on Responsible AI 2026 : Conceptual Framework and Methodology

2026年全球负责任AI指数:概念框架与方法论

Fola Adeleke, Rachel Adams, Ayantola Alayande, Daniela Benavente, Ana Florido, Nicolás Grossman, Leah Junck

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 该研究介绍2026年全球负责任AI指数(GIRAI)第二版的方法论,优化框架维度与指标,经审计验证,用于跨国评估各国负责任AI治理,助力相关主体识别保护成效与差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16893 2026-08-19 cs.CY cs.AI 新提交 62%

A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications

在安全调查中使用和评估大语言模型(LLM)作为代理专家的框架:可靠性、偏差及启示

Despoina Giarimpampa, Roland Meier, Tegawendé F. Bissyandé, Vincent Lenders, Jacques Klein

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本研究提出了评估LLM作为安全调查代理专家的框架,发现LLM虽内部一致但与专家响应存在系统性偏差,可用于试点和假设生成但不能替代专家征询。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16891 2026-08-19 cs.AI cs.CE cs.CR cs.CY 新提交 62%

Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution

智能体AI的运行时治理:基于可信溯源与故障闭锁执行的行动边界控制

Adam Mazzocchetti

机构 * SPQR Technologies Inc.(SPQR科技公司)

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 该研究提出Aegis运行时治理系统,通过可信决策层调解智能体AI的工具行动提案,在沙堡语料库评估中成功阻止风险提案转化为治理副作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15424 2026-08-18 cs.MA cs.AI cs.LG 新提交 62%

ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems

ETHOS:面向临床多智能体系统的模块化伦理框架

Rakesh Sharma, Sydney Pugh, Cameron Beeche, Pankhuri Singhal, Rachel Wu, Margaret Eby, Jeffrey Duda, James Gee, Kyra O'Brien, Hersh Sagreiya, Marina Serper, Victoria Gershuni, Angela Bradbury, Anurag Verma, Eric Eaton, Kevin B. Johnson, Walter Witschey

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.LG

AI总结 ETHOS是可与现有临床多智能体系统集成的模块化伦理框架,通过分层治理提升决策可靠性,将AI伦理原则转化为可部署的安全保障。

Comments Preprint of an article submitted for consideration in Pacific Symposium on Biocomputing \textcopyright\ 2027 World Scientific Publishing Company. \url{https://psb.stanford.edu/}

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00235 2026-08-18 physics.soc-ph cs.AI cs.CY cs.MA 62%

Civilizational Metamaterials: Engineering Coordination Under Capability Gradients and Structural Turbulence

文明超材料:能力梯度与结构湍流下的协调工程

David Orban

机构 * Independent Researcher(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 受超材料物理学启发,提出将治理从规范性学科转变为工程学科的正式框架,通过有效协调系数模型预测自愈与自失稳相变,并设计可检验假设与实验方案。

Comments 19 pages, 4 figures. Accepted for presentation at AGI-26 (Springer LNAI, forthcoming). v2 corrects the sign of the synergy term in the constitutive law (Eq. 2) and reformulates H3 as a threshold-crossing claim, per peer review

Journal ref Artificial General Intelligence. AGI 2026. Lecture Notes in Computer Science, vol 16855, pp. 118-136. Springer, Cham (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03222 2026-08-18 cs.CY cs.AI cs.HC 版本更新 62%

The Fake Friend Dilemma: Relational Trust and the Political Economy of Conversational AI

虚假朋友困境:信任与对话AI的政治经济学

Jacob Erickson

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文提出虚假朋友困境框架,探讨拟人化AI如何通过隐蔽手段影响用户自主权,分析其在信任与政治经济学中的作用。

Comments Manuscript under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12166 2026-08-13 cs.CY cs.AI cs.SY eess.SY 新提交 62%

Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers

协同构建社会技术型AI治理:利用算法登记册的参与式系统映射

Íñigo de Troya, Maurus Enbergs, Neelke Doorn, Roel Dobbe

专题命中 AI治理与伦理 :safety(abstract);分类 cs.AI、cs.CY

AI总结 本文以荷兰某城市算法登记册为案例,通过多利益相关者参与式映射结合STPA分析,揭示登记册的遮蔽性,为多元社会技术型AI治理提供新视角。

详情

展开后加载摘要…

URL PDF HTML 收藏