arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

潜在诊断分类学:构建分类器并诊断其决策的框架,应用于提示注入检测

The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection

Jaturong Kongmanee, Smile Thanapattheerakul

arXiv 2608.26423首次发表:更新:

发表机构

TrendAI™ Research(TrendAI™研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出潜在诊断分类学框架,构建提示注入检测分类器并诊断其可信决策,实验发现分类器约77%高置信度决策对单标记移除不鲁棒,分两种失败模式并给出修复策略。

AI 中文摘要

本文提出了一种框架,用于构建作为安全防护层的分类器,并开发互补诊断方法以识别分类器的哪些高置信度决策是可信的。该框架即潜在诊断分类学,包含三个部分:(i)构建维度优化的分类器,其嵌入维度通过交叉验证性能经验性选择,而非预先固定;(ii)定位相对少量的潜在支持向量(约占总训练样本的29%),这些样本代表用于识别改变分类器预测标签的标记的有影响力提示;(iii)利用此类标记及其相关攻击幅度构建诊断分类学。该诊断分类学为标记需不同处理的提示提供端到端指南:安全依赖分类器的决策;标记启发式偏差和启发式覆盖案例;将上下文不足案例路由以进行进一步人工/安全审查。将该框架应用于在公共提示注入数据集上训练的分类器,我们发现其高置信度决策的很大一部分(约77%)对移除单个标记不具有鲁棒性,且这种脆弱性分为两种不同的失败模式:置信度校准失败和真正可利用的捷径。对于分类学的每个区域,我们还推荐了诊断提示的修复策略。我们将该框架展示为一系列步骤,演示每一步的运作方式。

英文摘要

This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be trusted. This framework, the Latent Diagnostic Taxonomy, consists of (i) constructing a dimensionality-optimized classifier, in which the embedding dimensionality is empirically selected via cross-validated performance rather than fixed a priori, (ii) locating a relatively small set of latent support vectors (~ 29% of total training examples) representing influential prompts for identifying tokens that alter the classifier's predicted labels, and (iii) utilizing such tokens and their associated attack magnitudes for constructing a diagnostic taxonomy. This diagnostic taxonomy provides an end-to-end guideline for flagging prompts that require different treatments: rely Safely on the classifier's decision; flag Heuristic Bias and Heuristic Override cases; route Insufficient Context cases for further human/safety review. Applying the framework to a classifier trained on a public prompt injection dataset, we find that a substantial fraction of its confident decisions (~ 77%) are not robust to removing a single token, and that this brittleness separates into two distinct failure patterns: a confidence calibration failure and a genuinely exploitable shortcut. For each zone of the taxonomy, we also recommend strategies for remediating diagnosed prompts. We illustrate the framework as a series of steps, demonstrating how each step operates.

Comments10 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑