arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-11 至 2026-02-11 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 提示注入 2 篇

2601.09625 2026-02-11 cs.CR cs.AI 79%

The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multistep Malware Delivery Mechanism

提示武器杀链:提示注入如何逐渐演变成多步骤恶意软件交付机制

Oleg Brodt, Elad Feldman, Bruce Schneier, Ben Nassi

机构 * Department of Software and Information Systems Engineering, Ben-Gurion University of the Negev(本·古里安大学软件与信息系统工程系) School of Electrical and Computer Engineering, Tel Aviv University(特拉维夫大学电气与计算机工程学院) Harvard Kennedy School, Harvard University, and Munk School, University of Toronto(哈佛大学哈佛肯尼迪学校及多伦多大学穆克学校)

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

AI总结 本文提出提示武器杀链模型,揭示提示注入演变为多步骤恶意软件攻击机制的过程,并提出针对各阶段的防御策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09433 2026-02-11 cs.CR cs.AI 70%

Autonomous Action Runtime Management(AARM):A System Specification for Securing AI-Driven Actions at Runtime

自主动作运行时管理(AARM):为在运行时安全保护AI驱动的动作制定系统规范

Herman Errico

机构 * Independent Researcher, IEEE Member(独立研究者,IEEE会员)

专题命中 提示注入 :alignment(abstract);prompt injection(abstract);分类 cs.AI

AI总结 AARM提出了一种开放规范,旨在通过运行时安全机制保护AI驱动的动作,定义了拦截动作、评估策略和记录凭证等核心功能,以应对AI自主执行带来的安全挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏