arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14723cs.CR

检测与定位多源LLM智能体输入中的分段级投毒

Detecting and Localizing Segment-Level Poisoning in Multi-Source LLM-Agent Inputs

Xue Tan, Changhui Wang, Sanrui Yang, Hao Luan, Zhuyang Yu, Jin Wei, Ping Chen, Xiaoyan Sun, Jun Dai

首次发表
浏览论文内容

中文总结 AI 辅助

针对多源LLM智能体输入的分段级投毒,提出基于内部激活方向检测与定位的ActProbe框架,实现高效低开销的投毒检测与定位。

中文摘要 AI 辅助

现代大语言模型(LLM)智能体通常通过聚合来自多个外部来源的检索段落、用户评论和文档来构建提示。这种范式使其面临分段级投毒攻击,在这种攻击中,仅控制一小部分来源的对手会注入恶意内容以操纵模型输出。现有防御主要依赖文本模式、外部嵌入或辅助检测器,因此可能无法应对流畅且语义合理的投毒分段,并且在定位相关分段方面提供的支持有限。我们观察到,成功的证据破坏和对抗指令攻击会引发LLM内部激活的结构性偏移,形成一种一致的激活空间模式,我们称之为“投毒方向”。基于这一观察,我们提出了ActProbe,一个基于内部状态的框架,用于检测和定位多源LLM输入中的投毒分段。ActProbe将MLP激活投影到学习到的投毒方向上,并使用在小型校准集上训练的轻量级线性SVM来检测受污染的提示。随后,它应用BinRoL,该方法结合了递归替换消融、基于马氏距离的分支剪枝和基于MAD的稳健叶节点检测来定位投毒分段。ActProbe无需修改后端LLM,并将定位开销从O(n)的穷举探测降低到O(k log n)次前向传递。在三个数据集、两种攻击和四个开放权重LLM上,ActProbe实现了0.01的假阳性率、0.05的假阴性率、0.94的定位召回率和0.90的定位F1分数。它对抗防御感知的自适应攻击仍然有效,并可通过基于代理的投毒分段移除来保护黑盒API。

英文摘要

Modern large language model (LLM) agents often construct prompts by aggregating retrieved passages, user reviews, and documents from multiple external sources. This paradigm exposes them to segment-level poisoning attacks, in which an adversary controlling only a small subset of sources injects malicious content to manipulate model outputs. Existing defenses mainly rely on textual patterns, external embeddings, or auxiliary detectors and may therefore fail against fluent, semantically plausible poisoned segments. They also provide limited support for locating the responsible segments. We observe that successful corrupted-evidence and adversarial-instruction attacks induce structured shifts in the LLM's internal activations, forming a consistent activation-space pattern that we call the poison direction. Based on this observation, we propose ActProbe, an internal-state-based framework for detecting and localizing poisoned segments in multi-source LLM inputs. ActProbe projects MLP activations onto a learned poison direction and uses a lightweight linear SVM trained on a small calibration set to detect contaminated prompts. It then applies BinRoL, which combines recursive replacement ablation, Mahalanobis-distance-based branch pruning, and MAD-based robust leaf detection to locate poisoned segments. ActProbe requires no modification to the backend LLM and reduces localization overhead from O(n) exhaustive probing to O(k log n) forward passes. Across three datasets, two attacks, and four open-weight LLMs, ActProbe achieves a 0.01 false-positive rate, a 0.05 false-negative rate, 0.94 localization recall, and a 0.90 localization F1-score. It remains effective against defense-aware adaptive attacks and can protect black-box APIs through surrogate-based poisoned-segment removal.

发表机构

  • Fudan University(复旦大学)
  • Worcester Polytechnic Institute(伍斯特理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑