arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-24 至 2026-02-24 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 提示注入 2 篇

2602.13597 2026-02-24 cs.CR 86%

AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks

AlignSentinel: 一种考虑对齐的提示注入攻击检测方法

Yuqi Jia, Ruiqi Wang, Xilong Wang, Chong Xiang, Neil Gong

专题命中 提示注入 :prompt injection(title,abstract);alignment(title)

AI总结 AlignSentinel通过分析LLM注意力图,有效区分包含错位指令、对齐指令和非指令输入,提升提示注入攻击检测的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18514 2026-02-24 cs.CR cs.AI 85%

Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Models

招聘中的木马:针对标准与推理模型间接提示注入的红队案例研究

Manuel Wirth

机构 * University of Mannheim(曼海姆大学)

专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本研究通过红队测试揭示了标准与推理模型在间接提示注入中的安全差异,发现推理模型在复杂指令下易出现元认知泄漏,而标准模型在简单攻击中表现较弱。

Comments 43 pages, 3 synthetic CV PDF's, 6 chat history PDF's and system prompts. This work was developed as part of the Responsible AI course within the Mannheim Master in Data Science (MMDS) program at the University of Mannheim

详情

展开后加载摘要…

URL PDF HTML 收藏