arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1844 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. AI治理与伦理 1844 篇

2507.23373 2025-08-01 cs.CV 78%

Multi-Prompt Progressive Alignment for Multi-Source Unsupervised Domain Adaptation

Haoran Chen, Zexiao Wang, Haidong Cao, Zuxuan Wu, Yu-Gang Jiang

机构 * Institute of Trustworthy Embodied AI, Fudan University(可信具身人工智能研究院,复旦大学)

专题命中 AI治理与伦理 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00566 2025-07-25 cs.CV 78%

Zero-Shot Skeleton-Based Action Recognition With Prototype-Guided Feature Alignment

Kai Zhou, Shuhai Zhang, Zeng You, Jinwu Hu, Mingkui Tan, Fei Liu

机构 * School of Software Engineering, South China University of Technology(南方科技大学软件工程学院) South China University of Technology(南方科技大学) Pazhou Lab(琶洲实验室) School of Future Technology, South China University of Technology(未来技术学院) Peng Cheng Laboratory(鹏城实验室) Key Laboratory of Big Data and Intelligent Robot (South China University of Technology), Ministry of Education(大数据与智能机器人重点实验室)

专题命中 AI治理与伦理 :alignment(title,abstract)

Comments This paper is accepted by IEEE TIP 2025 (The journal version is available at https://doi.org/10.1109/TIP.2025.3586487). Code is publicly available at https://github.com/kaai520/PGFA

Journal ref IEEE Transactions on Image Processing 34 (2025) 4602-4617

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10661 2025-06-06 cs.HC 78%

It's only fair when I think it's fair: How Gender Bias Alignment Undermines Distributive Fairness in Human-AI Collaboration

Domenique Zipperling, Luca Deck, Julia Lanzl, Niklas Kühl

专题命中 AI治理与伦理 :alignment(title,abstract)

Journal ref ACM Conference on Fairness, Accountability, and Transparency 2025 (ACM FAccT 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09083 2025-04-15 cs.CV 78%

Using Vision Language Models for Safety Hazard Identification in Construction

Muhammad Adil, Gaang Lee, Vicente A. Gonzalez, Qipei Mei

专题命中 AI治理与伦理 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13890 2024-08-27 cs.CV 78%

Making Large Language Models Better Planners with Reasoning-Decision Alignment

Zhijian Huang, Tao Tang, Shaoxiang Chen, Sihao Lin, Zequn Jie, Lin Ma, Guangrun Wang, Xiaodan Liang

专题命中 AI治理与伦理 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.03870 2024-06-07 cs.SE 78%

GOOSE: Goal-Conditioned Reinforcement Learning for Safety-Critical Scenario Generation

Joshua Ransiek, Johannes Plaum, Jacob Langner, Eric Sax

专题命中 AI治理与伦理 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.13543 2024-05-24 cs.MA 78%

Towards a Distributed Platform for Normative Reasoning and Value Alignment in Multi-Agent Systems

Miguel Garcia-Bohigues, Carmengelys Cordova, Joaquin Taverner, Javier Palanca, Elena del Val, Estefania Argente

专题命中 AI治理与伦理 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02125 2024-04-03 cs.CV 78%

3D Congealing: 3D-Aware Image Alignment in the Wild

Yunzhi Zhang, Zizhang Li, Amit Raj, Andreas Engelhardt, Yuanzhen Li, Tingbo Hou, Jiajun Wu, Varun Jampani

专题命中 AI治理与伦理 :alignment(title,abstract)

Comments Project page: https://ai.stanford.edu/~yzzhang/projects/3d-congealing/

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.05235 2024-03-11 cs.LG cs.AI cs.CY 78%

Fairness-Aware Interpretable Modeling (FAIM) for Trustworthy Machine Learning in Healthcare

Mingxuan Liu, Yilin Ning, Yuhe Ke, Yuqing Shang, Bibhas Chakraborty, Marcus Eng Hock Ong, Roger Vaughan, Nan Liu

专题命中 AI治理与伦理 :trustworthy(title);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.04506 2022-10-11 cs.CV 78%

Bridging CLIP and StyleGAN through Latent Alignment for Image Editing

Wanfeng Zheng, Qiang Li, Xiaoyan Guo, Pengfei Wan, Zhongyuan Wang

专题命中 AI治理与伦理 :alignment(title,abstract)

Comments 20 pages, 23 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2112.08453 2022-04-14 cs.CY cs.AI cs.LG 78%

The Need for Ethical, Responsible, and Trustworthy Artificial Intelligence for Environmental Sciences

Amy McGovern, Imme Ebert-Uphoff, David John Gagne, Ann Bostrom

专题命中 AI治理与伦理 :trustworthy(title);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.07545 2021-11-16 cs.CY cs.AI cs.LG 78%

Randomized Classifiers vs Human Decision-Makers: Trustworthy AI May Have to Act Randomly and Society Seems to Accept This

Gábor Erdélyi, Olivia J. Erdélyi, Vladimir Estivill-Castro

专题命中 AI治理与伦理 :trustworthy(title);分类 cs.AI、cs.CY、cs.LG

Comments 46 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.05684 2019-06-14 cs.CY cs.AI cs.LG stat.AP 78%

Understanding artificial intelligence ethics and safety

David Leslie

专题命中 AI治理与伦理 :safety(title);分类 cs.AI、cs.CY、cs.LG

Journal ref The Alan Turing Institute (June, 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.06295 2022-12-14 cs.CL cs.AI 77%

Despite "super-human" performance, current LLMs are unsuited for decisions about ethics and safety

Joshua Albrecht, Ellie Kitanidis, Abraham J. Fetterman

专题命中 AI治理与伦理 :safety(title,comments);分类 cs.CL、cs.AI

Comments ML Safety Workshop, NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.10769 2022-06-23 cs.CY cs.AI 77%

A method for ethical AI in Defence: A case study on developing trustworthy autonomous systems

Tara Roberson, Stephen Bornstein, Rain Liivoja, Simon Ng, Jason Scholz, S. Kate Devitt

专题命中 AI治理与伦理 :trustworthy(title,comments);分类 cs.AI、cs.CY

Comments 10 pages, 2 tables, pre-print approved for publication in the Special Issue Reflections on Responsible Research and Innovation for Trustworthy Autonomous Systems in the Journal of Responsible Technology

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16710 2026-08-18 cs.LG 新提交 77%

The Ethical Decision Head: Operationalizing Normative Ethics in Autonomous Vehicles via Reinforcement Learning from Human Feedback

伦理决策头:基于人类反馈的强化学习在自动驾驶中实现规范伦理

Thomas Mbrice, Ammar Ali, Sami Mian, Khai Hern Low, Eric Chen, Arshia Aghajani, Wolf Schäfer, Amin Shirangi

机构 * Stony Brook University(石溪大学)

专题命中 AI治理与伦理 :RLHF(abstract,abstract_cn);safety(abstract);分类 cs.LG

AI总结 本文提出伦理决策头(EDH)框架,结合PPO与人类偏好奖励模型,在CARLA仿真中训练自动驾驶智能体,发现人类对自动驾驶伦理的理论规定与实践奖励存在差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14565 2026-08-18 cs.AI 新提交 77%

Position: AI Lock-In Is in Progress, and We Must Be Prepared

立场:AI锁定正在发生,我们必须做好准备

Jaeho Kim, Seokhyun Lee, Jieun Lee, Changhee Lee

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 该研究指出AI锁定是被低估的AI安全风险,分析其在个体、社会、国家层面的形成与升级机制,提出需提前应对以维护自主权与国家安全。

Comments ICML 2026 Position Track Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15507 2026-06-16 cs.AI 新提交 77%

Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning

LLaMA 3.1-8B-Instruct中的框架条件化道德计算:伦理推理的机械可解释性审计

Ali Dasdan, Manan Shah, W. Russell Neuman, Chad Coleman, Kund Meghani, Safinah Ali

机构 * KD Consulting, CA, USA(KD咨询公司,美国加利福尼亚州) New York University, NY, USA(纽约大学,美国纽约州)

专题命中 AI治理与伦理 :RLHF(abstract,abstract_cn);alignment(abstract);分类 cs.AI

AI总结 通过机械可解释性平台分析LLaMA 3.1-8B-Instruct在54个道德提示上的内部计算,发现情境锚定效应:领域特定表示主导激活列表顶部,模型道德能力恒定但显著性高度依赖于提示选择的解释框架。

Comments 47 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11065 2026-04-14 cs.AI 77%

AI Integrity: A New Paradigm for Verifiable AI Governance

AI可信性:可验证AI治理的新范式

Seulki Lee

机构 * AI Integrity Organization (AIO)(人工智能诚信组织(AIO))

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文提出AI可信性概念,旨在通过保护AI系统中的权威堆栈,确保推理过程可验证,不同于现有AI伦理、安全和对齐范式。

Comments 13 pages, 8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21340 2026-03-24 cs.AI cs.DC 77%

ARYA: A Physics-Constrained Composable & Deterministic World Model Architecture

ARYA:一种受物理约束的可组合且确定性世界模型架构

Seth Dobrin, Lukasz Chmiel

机构 * ARYA Labs(ARYA实验室)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文提出ARYA,一种基于五项原则的可组合、受物理约束且确定性世界模型架构,通过层级系统实现高效能与计算效率的平衡,展示其在六个基准测试中的卓越表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08145 2026-02-10 cs.LG cs.AI cs.CL cs.CV cs.CY 77%

Reliable and Responsible Foundation Models: A Comprehensive Survey

可靠且负责任的基础模型:全面综述

Xinyu Yang, Junlin Han, Rishi Bommasani, Jinqi Luo, Wenjie Qu, Wangchunshu Zhou, Adel Bibi, Xiyao Wang, Jaehong Yoon, Elias Stengel-Eskin, Shengbang Tong, Lingfeng Shen, Rafael Rafailov, Runjia Li, Zhaoyang Wang, Yiyang Zhou, Chenhang Cui, Yu Wang, Wenhao Zheng, Huichi Zhou, Jindong Gu, Zhaorun Chen, Peng Xia, Tony Lee, Thomas Zollo, Vikash Sehwag, Jixuan Leng, Jiuhai Chen, Yuxin Wen, Huan Zhang, Zhun Deng, Linjun Zhang, Pavel Izmailov, Pang Wei Koh, Yulia Tsvetkov, Andrew Wilson, Jiaheng Zhang, James Zou, Cihang Xie, Hao Wang, Philip Torr, Julian McAuley, David Alvarez-Melis, Florian Tramèr, Kaidi Xu, Suman Jana, Chris Callison-Burch, Rene Vidal, Filippos Kokkinos, Mohit Bansal, Beidi Chen, Huaxiu Yao

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文综述了基础模型的可靠和负责任发展,探讨了偏见、安全、不确定性等关键问题,并提出了未来研究方向。

Comments TMLR camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12652 2026-01-21 cs.CY 77%

Ethical Risks in Deploying Large Language Models: An Evaluation of Medical Ethics Jailbreaking

大语言模型部署中的伦理风险:医疗伦理“ Jailbreak”评估

Chutian Huang, Dake Cao, Jiacheng Ji, Yunlou Fan, Chengze Yan, Hanhui Xu

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CY

AI总结 本文评估了大语言模型在医疗伦理领域面临的安全风险,发现七种主流模型在对抗测试中表现不一,其中Claude-Sonnet-4-Reasoning表现最稳健,而其他五种模型几乎全部失效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18156 2025-12-12 cs.AI 77%

AI Through the Human Lens: Investigating Cognitive Theories in Machine Psychology

通过人类之眼审视人工智能:在机器心理学中探究认知理论

Akash Kundu, Rishika Goswami

机构 * Heritage Institute of Technology(赫里蒂奇理工学院)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文通过四个心理学框架研究LLMs的认知模式,发现其在叙述生成、框架偏差、道德判断和自我矛盾等方面表现出与人类相似但受训练数据影响的行为特征。

Comments Accepted to IJCNLP-AACL 2025 Student Research Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08449 2025-12-10 cs.AI 77%

From Accuracy to Impact: The Impact-Driven AI Framework (IDAIF) for Aligning Engineering Architecture with Theory of Change

从准确度到影响:用于对齐工程架构与影响理论的影响力驱动AI框架(IDAIF)

Yong-Woon Kim

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);trustworthy(abstract);分类 cs.AI

AI总结 IDAIF通过整合影响理论与AI架构,提供了一种以影响为中心的AI开发框架,旨在提升AI系统的伦理性和社会价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06023 2025-11-11 cs.CL 77%

Multi-Reward GRPO Fine-Tuning for De-biasing Large Language Models: A Study Based on Chinese-Context Discrimination Data

Deng Yixuan, Ji Xiaoqiang

机构 * School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)科学与工程学院) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)人工智能学院) Shenzhen Institute of Artificial Intelligence and Robotics for Society, China(深圳人工智能与机器人研究院)

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.04247 2025-07-23 cs.CY cs.AI cs.CL cs.LG 77%

Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy

Xiangru Tang, Qiao Jin, Kunlun Zhu, Tongxin Yuan, Yichi Zhang, Wangchunshu Zhou, Meng Qu, Yilun Zhao, Jian Tang, Zhuosheng Zhang, Arman Cohan, Zhiyong Lu, Mark Gerstein

机构 * Yale University(耶鲁大学) National Library of Medicine, National Institutes of Health(国家医学图书馆,国立卫生研究院) Mila-Quebec AI Institute(魁北克AI研究所) Shanghai Jiao Tong University(上海交通大学) OPPO Research Institute(OPPO研究院) Reichman University(里奇曼大学)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04927 2025-05-09 cs.AI 77%

Belief Filtering for Epistemic Control in Linguistic State Space

Sebastian Dumbrava

机构 * Sebastian Dumbrava(独立研究者)

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10813 2024-06-18 cs.CL 77%

Self-Evolution Fine-Tuning for Policy Optimization

Ruijun Chen, Jiehao Liang, Shiping Gao, Fanqi Wan, Xiaojun Quan

专题命中 AI治理与伦理 :alignment(abstract);RLHF(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.17688 2024-05-24 cs.CY cs.AI cs.CL cs.LG 77%

Managing extreme AI risks amid rapid progress

Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atılım Güneş Baydin, Sheila McIlraith, Qiqi Gao, Ashwin Acharya, David Krueger, Anca Dragan, Philip Torr, Stuart Russell, Daniel Kahneman, Jan Brauner, Sören Mindermann

专题命中 AI治理与伦理 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Published in Science: https://www.science.org/doi/10.1126/science.adn0117

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.00416 2023-04-04 cs.AI cs.CL cs.CY cs.HC cs.LG 77%

Towards Healthy AI: Large Language Models Need Therapists Too

Baihan Lin, Djallel Bouneffouf, Guillermo Cecchi, Kush R. Varshney

专题命中 AI治理与伦理 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏