arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-12-16 至 2025-12-16 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 提示注入 4 篇

2512.08290 2025-12-16 cs.CR cs.AI 85%

Systematization of Knowledge: Security and Safety in the Model Context Protocol Ecosystem

知识体系化:模型上下文协议生态系统中的安全与安全

Shiva Gaire, Srijan Gyawali, Saroj Mishra, Suman Niroula, Dilip Thakur, Umesh Yadav

机构 * Tribhuvan University(特里布文大学) University of North Dakota(北达科他大学) Youngstown State University(亚当斯州立大学) University of Missouri(密苏里大学) University of Toledo(托莱多大学)

专题命中 提示注入 :safety(title,abstract);alignment(abstract);prompt injection(abstract);分类 cs.AI

AI总结 本文系统化分析了模型上下文协议生态系统中的安全与安全风险,提出了全面的风险分类,并探讨了从对话聊天机器人到自主代理操作系统安全过渡的路线图。

Comments All authors contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12583 2025-12-16 cs.CR cs.AI 80%

Detecting Prompt Injection Attacks Against Application Using Classifiers

使用分类器检测应用于应用程序中的提示注入攻击

Safwan Shaheer, G. M. Refatul Islam, Mohammad Rafid Hamid, Md. Abrar Faiaz Khan, Md. Omar Faruk, Yaseen Nur

机构 * Department of Computer Science (CS)(计算机科学系) Department of Computer Science and Engineering (CSE)(计算机科学与工程系) School of Data and Sciences (SDS)(数据与科学学院)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

AI总结 本研究通过训练多种分类器,旨在检测网络应用中的提示注入攻击,提升系统安全性和稳定性。

Comments 9 pages, X figures; undergraduate research project on detecting prompt injection attacks against LLM integrated web applications using classical machine learning and neural classifiers

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23260 2025-12-16 cs.CR cs.AI 74%

From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows

从提示注入到协议利用:LLM驱动的AI代理工作流中的威胁

Mohamed Amine Ferrag, Norbert Tihanyi, Djallel Hamouda, Leandros Maglaras, Abderrahmane Lakas, Merouane Debbah

机构 * Department of Computer and Network Engineering, College of Information Technology, United Arab Emirates University(阿联酋大学计算机与网络工程系) Technology Innovation Institute(技术创新研究所) Eötvös Loránd University(埃奥瓦大学) Department of Computer Science, Guelma University(古尔马大学计算机科学系) De Montfort University(德蒙福特大学) Khalifa University of Science and Technology(卡里玛科学技术大学)

专题命中 提示注入 :prompt injection(title);分类 cs.AI

AI总结 本文提出了一种统一的端到端威胁模型,系统分类了LLM代理生态系统中的三十多种攻击技术,涵盖输入操纵、模型妥协、系统和隐私攻击及协议漏洞,并提供缓解策略以设计安全的代理AI系统。

Comments The paper is published in ICT Express (Elsevier)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12488 2025-12-16 cs.CL 57%

The American Ghost in the Machine: How language models align culturally and the effects of cultural prompting

机器中的美国幽灵:语言模型如何实现文化对齐以及文化提示的影响

James Luther, Donald Brown

机构 * School of Data Science University of Virginia(数据科学学院 多伦多大学)

专题命中 提示注入 :alignment(abstract);分类 cs.CL

AI总结 本研究通过文化提示测试了多种LLM对不同文化适应性,发现多数模型默认倾向美国,但对日本和中国文化对齐存在困难。

详情

展开后加载摘要…

URL PDF HTML 收藏