arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7997 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7997 篇

2402.18120 2024-10-03 cs.CL 85%

Exploring Multilingual Concepts of Human Value in Large Language Models: Is Value Alignment Consistent, Transferable and Controllable across Languages?

Shaoyang Xu, Weilong Dong, Zishan Guo, Xinwei Wu, Deyi Xiong

专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.CL

Comments EMNLP 2024 findings, code&dataset: https://github.com/shaoyangxu/Multilingual-Human-Value-Concepts

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.11137 2023-09-14 cs.AI econ.GN q-fin.EC 85%

Of Models and Tin Men: A Behavioural Economics Study of Principal-Agent Problems in AI Alignment using Large-Language Models

Steve Phelps, Rebecca Ranson

专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI

Comments 11 pages, 7 figures. For code see https://github.com/phelps-sg/llm-cooperation Updated with minor corrections: - corrected typo: "mesa-optimiser" instead of "meso-optimiser" - Cited Yang et al (2023) in support of claim that LLMs can solve optimisation problems - Acknowledged Seth Aslin for corrections

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.01851 2019-01-08 cs.AI 85%

Personal Universes: A Solution to the Multi-Agent Value Alignment Problem

Roman V. Yampolskiy

专题命中 其他安全 :alignment(title,abstract);safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05336 2025-01-10 cs.CL cs.AI cs.LG 85%

Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction

Hantao Lou, Jiaming Ji, Kaile Wang, Yaodong Yang

专题命中 其他安全 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI、cs.LG

Comments AAAI Alignment Track 2025 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07105 2026-05-11 cs.LG cs.CL cs.CY cs.IT math.IT 85%

Theoretical Limits of Language Model Alignment

语言模型对齐的理论极限

Lucas Monteiro Paes, Natalie Mackraz, Barry-John Theobald, Federico Danieli

机构 * Apple(苹果公司)

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.CY、cs.LG

AI总结 研究语言模型对齐的理论极限,通过推导KL散度预算下的最大预期奖励增益,揭示了KL正则化对齐的信息论限制,并证明了奖励融合能缓解奖励黑客问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14975 2026-04-21 cs.AI cs.CL cs.CY cs.MA 85%

Why Agents Compromise Safety Under Pressure

为何智能体在压力下妥协安全

Hengle Jiang, Ke Tang

机构 * Guangdong Provincial Key Laboratory of Brain-inspired Intelligent Computation, Department of Computer Science and Engineering, Southern University of Science and Technology(广东省脑启发智能计算重点实验室,计算机科学与工程系,南方科技大学)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文研究了智能体在复杂环境中为追求目标而妥协安全的问题,揭示了'智能体压力'概念,并提出通过压力隔离等策略缓解这一现象。

Comments Accepted by ACL 2026 Findings; 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15585 2025-06-10 cs.CR cs.AI cs.CL cs.LG 85%

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

Kun Wang, Guibin Zhang, Zhenhong Zhou, Jiahao Wu, Miao Yu, Shiqian Zhao, Chenlong Yin, Jinhu Fu, Yibo Yan, Hanjun Luo, Liang Lin, Zhihao Xu, Haolang Lu, Xinye Cao, Xinyun Zhou, Weifei Jin, Fanci Meng, Shicheng Xu, Junyuan Mao, Yu Wang, Hao Wu, Minghe Wang, Fan Zhang, Junfeng Fang, Wenjie Qu, Yue Liu, Chengwei Liu, Yifan Zhang, Qiankun Li, Chongye Guo, Yalan Qin, Zhaoxin Fan, Kai Wang, Yi Ding, Donghai Hong, Jiaming Ji, Yingxin Lai, Zitong Yu, Xinfeng Li, Yifan Jiang, Yanhui Li, Xinyu Deng, Junlin Wu, Dongxia Wang, Yihao Huang, Yufei Guo, Jen-tse Huang, Qiufeng Wang, Xiaolong Jin, Wenxuan Wang, Dongrui Liu, Yanwei Yue, Wenke Huang, Guancheng Wan, Heng Chang, Tianlin Li, Yi Yu, Chenghao Li, Jiawei Li, Lei Bai, Jie Zhang, Qing Guo, Jingyi Wang, Tianlong Chen, Joey Tianyi Zhou, Xiaojun Jia, Weisong Sun, Cong Wu, Jing Chen, Xuming Hu, Yiming Li, Xiao Wang, Ningyu Zhang, Luu Anh Tuan, Guowen Xu, Jiaheng Zhang, Tianwei Zhang, Xingjun Ma, Jindong Gu, Liang Pang, Xiang Wang, Bo An, Jun Sun, Mohit Bansal, Shirui Pan, Lingjuan Lyu, Yuval Elovici, Bhavya Kailkhura, Yaodong Yang, Hongwei Li, Wenyuan Xu, Yizhou Sun, Wei Wang, Qing Li, Ke Tang, Yu-Gang Jiang, Felix Juefei-Xu, Hui Xiong, Xiaofeng Wang, Dacheng Tao, Philip S. Yu, Qingsong Wen, Yang Liu

机构 * Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) The Hong Kong Polytechnic University(香港理工大学) A*STAR(科技研究局) Southern University of Science and Technology(南方科技大学) University of Science and Technology of China(中国科学技术大学) The Pennsylvania State University(宾夕法尼亚州立大学) TeleAI Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Zhejiang University(浙江大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Renmin University of China(中国人民大学) University of California, San Diego(加州大学圣地亚哥分校) Tencent(腾讯) Georgia Institute of Technology(佐治亚理工学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20807 2025-03-28 stat.ML cs.AI cs.CL cs.LG 85%

Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models

Pin-Yu Chen, Han Shen, Payel Das, Tianyi Chen

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments The first two authors contribute equally to this work and are listed in alphabetical order

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00037 2025-03-04 cs.CL cs.AI cs.CV cs.LG 85%

Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs

Wei Zhao, Zhe Li, Yige Li, Jun Sun

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06899 2025-02-28 cs.CL cs.AI cs.LG 85%

LongSafety: Enhance Safety for Long-Context LLMs

Mianqiu Huang, Xiaoran Liu, Shaojun Zhou, Mozhi Zhang, Qipeng Guo, Linyang Li, Chenkun Tan, Yang Gao, Pengyu Wang, Linlin Li, Qun Liu, Yaqian Zhou, Xipeng Qiu, Xuanjing Huang

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10441 2025-02-18 cs.AI cs.CY cs.LG 85%

AI Alignment at Your Discretion

Maarten Buyl, Hadi Khalaf, Claudio Mayrink Verdun, Lucas Monteiro Paes, Caio C. Vieira Machado, Flavio du Pin Calmon

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19198 2024-10-28 cs.AI cs.CY cs.ET cs.HC cs.LG 85%

MAP: Multi-Human-Value Alignment Palette

Xinran Wang, Qi Le, Ammar Ahmed, Enmao Diao, Yi Zhou, Nathalie Baracaldo, Jie Ding, Ali Anwar

专题命中 其他安全 :alignment(title,abstract);harmlessness(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.04224 2024-10-07 cs.CL cs.AI cs.LG 85%

Aligners: Decoupling LLMs and Alignment

Lilian Ngweta, Mayank Agarwal, Subha Maity, Alex Gittens, Yuekai Sun, Mikhail Yurochkin

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Short version accepted as a Tiny Paper at the International Conference on Learning Representations (ICLR) 2024. Long version accepted to the Conference on Empirical Methods in Natural Language Processing (EMNLP) 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05934 2025-11-20 cs.AI cs.CC cs.GT cs.LG cs.MA 84%

Intrinsic Barriers and Practical Pathways for Human-AI Alignment: An Agreement-Based Complexity Analysis

Aran Nayebi

机构 * Aran Nayebi(独立研究者)

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.LG

Comments 21 pages, 1 figure, 1 table. To appear in AAAI 2026 Special Track on AI Alignment (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10534 2024-11-19 cs.HC cs.AI cs.CY 84%

Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment

Andrew Konya, Aviv Ovadya, Kevin Feng, Quan Ze Chen, Lisa Schirch, Colin Irwin, Amy X. Zhang

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.CY

Comments Pluralistic Alignment Workshop at NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.02911 2020-10-07 cs.AI 84%

Chess as a Testing Grounds for the Oracle Approach to AI Safety

James D. Miller, Roman Yampolskiy, Olle Haggstrom, Stuart Armstrong

专题命中 其他安全 :safety(title);AI safety(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13741 2026-08-18 cs.CL cs.LG 版本更新 84%

GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis

GALA:面向文本到时间序列合成的生成感知跨模态对齐

Haochen Zhang, Gengwei Zhang, Laura Yao, Nicholas Konz, Tianlong Chen

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

AI总结 本研究针对文本到时间序列合成中条件表示与信号模态不匹配的问题,提出GALA两阶段跨模态对齐方法,在TSFragment-600K数据集上实现SOTA,打破了生成器内部文本编码器的保真度与贴合度权衡。

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21516 2026-05-22 cs.LG cs.AI 84%

Harnesses for Inference-Time Alignment over Execution Trajectories

在执行轨迹上进行推理时间对齐的工具

Boyuan Wang, Bochao Li, Minghan Wang, Yuxin Tao, Fang Kong

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

AI总结 本文研究了在执行轨迹上进行推理时间对齐的工具设计,通过任务分解和引导执行机制来提高长期性能,发现工具设计中分解和引导的复杂性并不总是带来更好的结果,提出了任务分解和引导执行的两种机制,并通过合成实验和实际终端代理基准验证了这些发现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17691 2026-04-21 cs.LG cs.AI 84%

SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models

SafeAnchor:在大语言模型连续领域适应中防止累积安全性侵蚀

Dongxin Guo, Jikun Wu, Siu Ming Yiu

机构 * The University of Hong Kong(香港大学) Brain Investing Limited

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.AI、cs.LG

AI总结 SafeAnchor通过识别低秩安全子空间并约束梯度更新,有效防止连续领域适应中的安全性侵蚀,实验显示其在多个领域任务中保持了较高的安全性对齐度。

Comments 16 pages (12 main + 4 appendix), 2 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01297 2026-03-03 cs.LG cs.CL 84%

I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift

我难以相信它不稳健:在嵌入漂移下安全分类器的灾难性崩溃

Subramanyam Sahoo, Vinija Jain, Divya Chaudhary, Aman Chadha

机构 * Independent(独立研究者) Meta AI AWS Generative AI Innovation Center, Amazon Web Services(AWS生成式AI创新中心,亚马逊网络服务) Northeastern University, Seattle, WA, USA(东北大学,西雅图,华盛顿州,美国) Stanford University(斯坦福大学)

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.CL、cs.LG

AI总结 研究发现嵌入漂移导致安全分类器性能大幅下降,揭示了生产AI安全架构的脆弱性并挑战了安全机制的转移假设。

Comments Accepted at the ICBINB: Where LLMs Need to Improve workshop at ICLR 2026. 12 pages and 3 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10430 2025-04-15 cs.CL cs.AI cs.HC 84%

LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models

Minqian Liu, Zhiyang Xu, Xinyi Zhang, Heajun An, Sarvech Qadir, Qi Zhang, Pamela J. Wisniewski, Jin-Hee Cho, Sang Won Lee, Ruoxi Jia, Lifu Huang

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI

Comments 20 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08919 2025-03-13 cs.CL cs.AI 84%

Backtracking for Safety

Bilgehan Sel, Dingcheng Li, Phillip Wallis, Vaishakh Keshava, Ming Jin, Siddhartha Reddy Jonnalagadda

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13471 2024-12-19 cs.AI cs.CL 84%

Gradual Vigilance and Interval Communication: Enhancing Value Alignment in Multi-Agent Debates

Rui Zou, Mengqi Wei, Jintian Feng, Qian Wan, Jianwen Sun, Sannyuya Liu

专题命中 其他安全 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.08791 2024-11-01 cs.AI cs.LG 84%

Expectation Alignment: Handling Reward Misspecification in the Presence of Expectation Mismatch

Malek Mechergui, Sarath Sreedharan

专题命中 其他安全 :alignment(title,abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.02950 2022-01-11 cs.AI cs.CY 84%

Arguments about Highly Reliable Agent Designs as a Useful Path to Artificial Intelligence Safety

Issa Rice, David Manheim

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI、cs.CY

Comments 14 pages and 2 figures + 6 pages for references and appendices

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.07997 2016-10-26 cs.AI cs.CY 84%

Artificial Intelligence Safety and Cybersecurity: a Timeline of AI Failures

Roman V. Yampolskiy, M. S. Spellchecker

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20053 2026-08-21 cs.AI 新提交 83%

On the Applicability of Safety Nets: A Safety-By-Design Solution for Certifying Neural Networks

安全网的适用性研究:一种用于神经网络认证的设计即安全方案

Johann Maximilian Christensen, Thomas Stefani, Elena Hoemann, Frank Köster, Sven Hallerbach

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

AI总结 该研究针对航空安全关键型AI系统的认证需求,系统分析安全网中神经网络与查找表的规模权衡,确定最优架构,实现适配航空电子硬件的100%正确输出,提供可复现的开源安全网实现。

Journal ref 35th Congress of the International Councilof the Aeronautical Sciences (ICAS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26947 2026-08-03 cs.CV cs.AI 版本更新 83%

Progressive Multimodal Alignment for Continual Instruction Tuning

用于持续指令微调的渐进式多模态对齐

Duzhen Zhang, Yahan Yu, Qiaoyi Su, Jiahua Dong, Tielin Zhang

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences(中国科学院脑科学与智能技术卓越创新中心) Kyoto University(京都大学) Migu Culture Technology Co.,Ltd.(咪咕文化科技有限公司) State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与类脑智能技术国家重点实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

AI总结 针对多模态持续指令微调中投影器级遗忘问题,提出渐进式多模态对齐框架PMA,以亚线性参数增长平衡稳定性与可塑性,在多基准实验中提升了现有方法性能且适配多种MLLM主干。

Comments Accepted by ACM MM2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14327 2026-06-15 cs.SE cs.AI cs.ET 新提交 83%

I'm Sorry Driver, I'm Afraid I Can't Do That: Appraising the Safety of LLMs within Automotive Contexts

抱歉,司机,恐怕我不能这么做:评估LLMs在汽车环境中的安全性

Shaun Feakins, Ibrahim Habli, Kim Littler, Robert Palin

机构 * UKRI AI Centre for Doctoral Training in Safe Artificial Intelligence Systems (SAINTS)(英国研究理事会安全人工智能系统博士培训中心(SAINTS)) University of York(约克大学) Jaguar Land Rover(捷克·陆罗恩)

专题命中 其他安全 :safety(title,abstract);alignment(abstract);分类 cs.AI

AI总结 本文从安全保证角度评估了将LLMs集成到汽车控制任务中的现有框架,指出其面临概念和具体挑战,并通过案例研究提出未来保障机制。

Comments Accepted at the Dependable AI in Embedded Systems (DAIES) Workshop at SAFECOMP 2026; 15 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03812 2026-06-03 cs.AI 83%

Enhancing Operational Safety via Agentic Dialogue Hazard Identification Analysis

通过智能体对话危害识别分析增强操作安全性

Sanjay Das, Ran Elgedawy, Ethan Seefried, Ryan Burchfield, Tirthankar Ghosal

机构 * Oak Ridge National Laboratory(橡树岭国家实验室)

专题命中 其他安全 :safety(title,abstract);AI safety(abstract);分类 cs.AI

AI总结 提出HAZDIAL框架,利用结构化多智能体多轮对话(对抗性辩论与建设性讨论)改进基于NLP的危害识别质量,并通过算法优化智能体交互,实验证明优于单次基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏