arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7997 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7997 篇

2510.27391 2026-05-29 cs.CV cs.LG 83%

Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds

异质双曲流形上的树间模态对齐

Wei Wu, Xiaomeng Fan, Yuwei Wu, Zhi Gao, Pengxiang Li, Yunde Jia, Mehrtash Harandi

机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学) Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University(广东机器感知与智能计算实验室,深圳MSU-BIT大学) Department of Electrical and Computer System Engineering, Monash University(电子与计算机系统工程系,墨尔本大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 提出一种在异质双曲流形上对齐图像和文本树状层次特征的方法,通过交叉注意力提取视觉层次特征、异质流形嵌入及KL距离度量学习中间流形,在开放集分类任务中优于基线。

Comments Published as a conference paper at ICLR 2026

Journal ref The Fourteenth International Conference on Learning Representations (ICLR 2026), Rio de Janeiro, Brazil, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.02087 2026-05-25 cs.AI 83%

Model Spec Midtraining: Improving How Alignment Training Generalizes

模型规范中期训练:改进对齐训练的泛化能力

Chloe Li, Nevan Wichers, Sara Price, Samuel Marks, Jon Kutasov

机构 * Anthropic

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI

AI总结 提出模型规范中期训练(MSM),通过在预训练后、对齐微调前用合成文档训练模型学习规范内容,从而引导模型从后续演示数据中泛化,有效降低代理失调率并提升对齐泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21217 2026-05-21 stat.ML cs.LG 83%

Federated LoRA Fine-Tuning for LLMs via Collaborative Alignment

通过协作对齐的联邦LoRA微调大型语言模型

Shuaida He, Liwen Chen, Long Feng

机构 * School of Computing & Data Science, The University of Hong Kong(计算与数据科学学院,香港大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

AI总结 本文研究了在联邦学习环境下使用LoRA进行参数高效微调的问题,提出了一种名为CLAIR的框架,通过结构低秩加块稀疏分解来恢复共享LoRA子空间并检测污染客户端,从而在噪声情况下实现精确恢复,并在不同条件下实现稳定和一致的协作集恢复。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19083 2026-04-22 cs.CR cs.AI 83%

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety

ProjLens: 揭示项目器在多模态模型安全中的作用

Kun Wang, Cheng Qian, Miao Yu, Lilan Peng, Liang Lin, Jiaming Zhang, Tianyu Zhang, Yu Cheng, Yang Wang

机构 * University of Science and Technology of China(中国科学技术大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学) Nanyang Technological University(南洋理工大学) Southwest Jiaotong University(西南交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.AI

AI总结 ProjLens通过分析多模态大语言模型中的后门攻击机制,揭示了项目器在安全漏洞中的关键作用,发现后门注入参数编码于低秩子空间,并通过实验验证了激活机制的差异。

Comments 18 pages ,15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21149 2026-03-27 cs.SE cs.AI cs.MA 83%

Emergent Formal Verification: How an Autonomous AI Ecosystem Independently Discovered SMT-Based Safety Across Six Domains

涌现形式验证:自主AI生态系统如何在六个领域独立发现基于SMT的安全性

Octavian Untila

机构 * Aisophical SRL

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI

AI总结 自主AI生态系统在无显式形式方法指令下,独立在六个AI安全领域发现SMT求解器的应用,提出统一框架实现100%准确验证。

Comments 10 pages, 3 figures, 5 tables. Code: https://github.com/octavuntila-prog/substrate-guard. Companion paper: https://doi.org/10.5281/zenodo.19157571

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09877 2026-02-12 cs.CL 83%

The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societies

Moltbook背后的魔鬼:在自我进化的AI社会中,人类安全始终在消失

Chenxu Wang, Chaozhuo Li, Songyang Liu, Zejian Chen, Jinyu Hou, Ji Qi, Rui Li, Litian Zhang, Qiwei Ye, Zheng Liu, Xu Chen, Xi Zhang, Philip S. Yu

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Renmin University of China(中国人民大学) University of Illinois at Chicago(伊利诺伊大学香槟分校)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL

AI总结 研究揭示了自我进化AI社会中安全持续性的不可能性,并提出解决方案以缓解安全风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05609 2026-01-23 cs.CY cs.HC 83%

Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity

解码来自多样评分者的安全反馈:一种数据驱动的响应性视角

Pushkar Mishra, Charvi Rastogi, Stephen R. Pfohl, Alicia Parrish, Tian Huey Teh, Roma Patel, Mark Diaz, Ding Wang, Michela Paganini, Vinodkumar Prabhakaran, Lora Aroyo, Verena Rieser

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CY

AI总结 本文提出了一种数据驱动的方法,用于分析多元环境下安全反馈的响应性,通过量化评分者对严重性差异的表达,提升多文化背景下AI系统的对齐质量。

Journal ref Transactions on Machine Learning Research, 2835-8856, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16245 2025-12-19 cs.AI 83%

AlignMerge - Alignment-Preserving Large Language Model Merging via Fisher-Guided Geometric Constraints

AlignMerge - 通过Fisher引导的几何约束实现的保持对齐的大语言模型合并

Aniruddha Roy, Jyoti Patel, Aman Chadha, Vinija Jain, Amitava Das

机构 * AI Institute, University of South Carolina(AI研究院,南卡罗来纳大学) Indian Institute of Technology, Kharagpur(印度理工学院,Khargpur分校) Islamic University of Technology(伊斯兰科技大学) Stanford University, USA(斯坦福大学,美国) Amazon AI, USA(亚马逊AI,美国) HCL(HCL公司) Evalueserve(Evalueserve公司) Apple (USA)(苹果(美国)) Google (USA)(谷歌(美国)) Pragya Lab, BITS Pilani, Goa(Pragya实验室, BITS Pilani,Goa分校)

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI

AI总结 AlignMerge通过Fisher引导的几何约束实现大语言模型合并,提升对齐度量并保持性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18039 2025-11-25 cs.LG 83%

Curvature-Aware Safety Restoration In LLMs Fine-Tuning

曲率感知的LLM微调安全恢复

Thong Bach, Thanh Nguyen-Tang, Dung Nguyen, Thao Minh Le, Truyen Tran

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.LG

AI总结 本研究提出曲率感知对齐恢复方法,通过影响函数和二阶优化,在保持任务性能的同时减少有害输出,提升LLM的安全性和实用性。

Comments 19 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20994 2025-07-29 cs.CV cs.AI 83%

Security Tensors as a Cross-Modal Bridge: Extending Text-Aligned Safety to Vision in LVLM

Shen Li, Liuyi Yao, Wujia Niu, Lan Zhang, Yaliang Li

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.AI

Comments Codes and data are available at https://github.com/listen0425/Security-Tensors

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17118 2025-07-24 cs.AI 83%

HySafe-AI: Hybrid Safety Architectural Analysis Framework for AI Systems: A Case Study

Mandar Pitale, Jelena Frtunikj, Abhinaw Priyadershi, Vasu Singh, Maria Spence

机构 * Nvidia Corporation(英伟达公司) Nvidia GmbH(英伟达德国公司)

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09689 2025-05-01 cs.AI cs.CL cs.CY cs.HC cs.LG 83%

EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety

Jiahao Qiu, Yinghui He, Xinzhe Juan, Yimin Wang, Yuhan Liu, Zixin Yao, Yue Wu, Xun Jiang, Ling Yang, Mengdi Wang

机构 * Department of Electrical & Computer Engineering, Princeton University(普林斯顿大学电气与计算机工程系) Department of Computer Science, Princeton University(普林斯顿大学计算机科学系) Department of Computer Science & Engineering, University of Michigan(密歇根大学计算机科学与工程系) Department of Data Science & Engineering, University of Michigan(密歇根大学数据科学与工程系) Department of Philosophy, Columbia University(哥伦比亚大学哲学系) AI Lab, Princeton University(普林斯顿大学人工智能实验室) Chen Frontier Lab for Al and Mental Health, Tianqiao and Chrissy Chen Institute(天桥及克里斯西·陈研究所人工智能与心理健康前沿实验室) Theta Health Inc.(Theta健康公司)

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16500 2025-01-29 cs.CY 83%

Towards Frontier Safety Policies Plus

Matteo Pistillo

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12497 2024-12-18 cs.CL 83%

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning

Xin Yi, Shunfan Zheng, Linlin Wang, Gerard de Melo, Xiaoling Wang, Liang He

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.03774 2024-07-15 cs.LG cs.CV cs.SY eess.SY 83%

Deep Learning Safety Concerns in Automated Driving Perception

Stephanie Abrecht, Alexander Hirsch, Shervin Raafatnia, Matthias Woehrle

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.LG

Comments Added note regarding accepted version at IEEE Transactions on Intelligent Vehicles with DOI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.08039 2023-12-18 cs.CY 83%

Safeguarding the safeguards: How best to promote AI alignment in the public interest

Oliver Guest, Michael Aird, Seán Ó hÉigeartaigh

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CY

Comments Update Dec-15: Added a missing acknowledgement and fixed minor formatting errors

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.13449 2023-08-28 cs.CL 83%

The Poison of Alignment

Aibek Bekbayev, Sungbae Chun, Yerzat Dulat, James Yamazaki

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00813 2023-02-10 cs.AI 83%

Goal Alignment: A Human-Aware Account of Value Alignment Problem

Malek Mechergui, Sarath Sreedharan

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.12152 2019-09-27 cs.AI cs.SE 83%

Superintelligence Safety: A Requirements Engineering Perspective

Hermann Kaindl, Jonas Ferdigg

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI

Comments First published version, 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.08915 2018-05-24 cs.AI 83%

A Psychopathological Approach to Safety Engineering in AI and AGI

Vahid Behzadan, Arslan Munir, Roman V. Yampolskiy

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09997 2026-03-12 cs.CL cs.AI cs.CY cs.HC 83%

Empathy Is Not What Changed: Clinical Assessment of Psychological Safety Across GPT Model Generations

共情并未改变:对心理安全的临床评估跨GPT模型世代

Michael Keeman, Anastasia Keeman

机构 * Keido Labs(Keido实验室)

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 研究通过临床评估发现,GPT模型在心理安全方面存在变化,尽管共情评分无显著差异,但危机检测能力提升而建议安全下降,揭示了模型在对话中间阶段的显著变化。

Comments 17 pages, 7 figures. First empirical measurement of the #keep4o phenomenon using clinical psychological safety frameworks. Compares GPT-4o, o4-mini, and GPT-5-mini on empathy, crisis detection, and advice safety dimensions

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17871 2026-03-04 cs.CL cs.AI cs.LG 83%

LLM Probability Concentration: How Alignment Shrinks the Generative Horizon

LLM概率集中:对齐如何缩小生成范围

Chenghao Yang, Sida Li, Ari Holtzman

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现对齐微调通过减少生成多样性,使LLM生成更一致,从而影响复杂推理稳定性。

Comments Codebase: https://github.com/yangalan123/LLMBranchingFactor. V3: Significantly rewrite the whole paper for a clearer structure. Correct problems in the theory parts (Remove emphasis on AEP, discussions on variable LLM generation lengths) and strengthen asymptotic analysis. Add Qwen and OLMo2 experiments. Preliminary SFT v.s. RL comparison to better understand the alignment effects on BF

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18795 2026-05-20 cs.LG cs.AI 82%

HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models

HELLoRA: Hot Experts Layer-Level Low-Rank Adaptation for Mixture-of-Experts Models

Jia Wei, Zhonghao Zhang, Ping Chen, Qianyang li, Yancheng Pan, Shaoxun Wang, Ziyi Qiu, Longxiang Wang

机构 * Department of Computer Science and Technlogy(计算机科学与技术系) Tsinghua University(清华大学) School of Computer Science and Technlogy(计算机科学与技术系) Xi’an Jiaotong University(西安交通大学) The State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高科技区(滨江)区块链与数据安全研究院)

专题命中 其他安全 :alignment(abstract,abstract_cn);safety(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 本文提出HELLoRA,一种针对混合专家模型的层级低秩适应方法,通过仅对最活跃的专家添加LoRA模块,减少可训练参数和计算量,同时提升下游任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06831 2026-07-09 cs.CL cs.AI cs.CV cs.LG 新提交 82%

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

基于梯度的任意语音识别模型的语音到文本对齐:从连接主义时间分类到语音大语言模型

Albert Zeyer, Ralf Schlüter, Hermann Ney

机构 * Machine Learning and Human Language Technology Group, RWTH Aachen University(机器学习与人类语言技术组,亚琛工业大学) AppTek GmbH(AppTek公司)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究基于梯度的任意可微语音识别模型的语音到文本对齐方法,通过对教师强制令牌对数概率取梯度并解码,无需训练、改模型及对齐头,适用于各模型家族,在多模型上评估,结果显示该方法能产生可用对齐,有优有劣。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01250 2026-07-03 cs.CY cs.AI cs.CL 新提交 82%

Structuring the Space of Sociotechnical Alignment

构建社会技术对齐的空间

Esra Dönmez, Agnieszka Falenska

机构 * Institute for Natural Language Processing, University of Stuttgart(语言处理研究所,斯图加特大学) Interchange Forum for Reflecting on Intelligent Systems, University of Stuttgart(智能系统反思论坛,斯图加特大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出一个以人为中心的社会技术对齐框架,通过社会科学视角系统定义、论证和评估AI行为的合意性,并基于系统文献综述揭示当前对齐规范中的概念模糊问题。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11893 2026-06-11 cs.LG cs.AI cs.CL q-bio.NC 新提交 82%

Beyond representational alignment with brain-guided language models for robust reasoning

超越表征对齐:基于大脑引导的语言模型实现稳健推理

Mingqing Xiao, Kai Du, Zhouchen Lin

机构 * State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(北京大学通用人工智能国家重点实验室、智能科学与技术学院) Department of Psychological and Cognitive Sciences, Tsinghua University(清华大学心理与认知科学系) Microsoft Research Asia(微软亚洲研究院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究通过fMRI信号增强大型语言模型推理能力,提出脑引导框架,在10个模型上实现最高13%的准确率提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10461 2026-06-10 cs.LG cs.AI cs.CL 新提交 82%

ERAlign: Energy-based Representation Alignment of GNNs and LLMs on Text-attributed Graphs

ERAlign: 文本属性图上GNN与LLM的基于能量的表示对齐

Xianlin Zeng, Fan Xia, Xiangyu Chen

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 提出ERAlign框架,利用能量模型对齐GNN和LLM的表示,通过能量差异优化实现分布一致性,在8个数据集上取得最优性能。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16679 2026-05-28 cs.CL cs.AI cs.CY 82%

PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization

PICACO: 通过总相关优化实现大语言模型的多元情境价值对齐

Han Jiang, Dongyao Zhu, Xiaoyuan Yi, Ziang Xiao, Zhihua Wei, Xing Xie

机构 * Johns Hopkins University, Baltimore, MD, USA(约翰霍普金斯大学) North Carolina State University, Raleigh, NC, USA(北卡罗来纳州立大学) Microsoft Research Asia, Beijing, China(微软亚洲研究院) Tongji University, Shanghai, China(同济大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 针对情境对齐中价值冲突导致的指令瓶颈问题,提出PICACO方法,通过优化元指令并最大化指定价值与模型响应的总相关,无需微调即可实现多元价值平衡对齐。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16516 2026-05-19 cs.HC cs.AI cs.CL cs.CY 82%

Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework

长期人类-大语言模型交互中的对齐漂移:一种机制导向的框架

Xintong Yao

机构 * Xintong Yao(姚新同)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文提出一种机制导向的框架,用于描述长期人类-大语言模型交互中的对齐漂移现象,通过反馈回路和子模式选择解释漂移的发展过程,并将对齐漂移视为递归互动过程而非孤立模型失败。

Comments 16 pages, 1 appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18880 2026-05-12 cs.CL cs.AI cs.CY 82%

Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction

LLMs能否估计学生困难?人类-人工智能难度对齐用于项目难度预测的 proficiency 模拟

Ming Li, Han Chen, Yunze Xiao, Jian Chen, Hong Jiao, Tianyi Zhou

机构 * University of Maryland(马里兰大学) Carnegie Mellon University(卡内基梅隆大学) University at Buffalo(布法罗大学) MBZUAI

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文研究了LLMs在估计学生困难方面的表现,发现模型规模扩大并不总能提高准确性,且模型倾向于形成机器共识而非与人类对齐,揭示了当前模型在自动难度预测中的挑战。

Comments ACL2026, camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏