arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13534cs.CL

有害性传播动力学:大型语言模型中对抗意图的逐层轨迹

Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models

Noor Islam S. Mohammad, Uluğ Bayazıt

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出有害性传播动力学(HPD),发现有害提示的隐藏状态投影随层数单调上升,并据此构建轻量级审核器Herald,以288参数MLP分类七维特征,在多个基准上超越现有防护模型,实现高效且可解释的对抗意图检测。

中文摘要 AI 辅助

我们识别出\u201c有害性传播动力学(HPD)\u201d:对于有害提示,最后一个token的隐藏状态在习得的有害方向上的投影随Transformer深度单调上升,而良性提示则保持平坦或振荡。这种跨层特征将有害意图反映为一种\u201c逐步解析\u201d的语义属性:表面形式早期出现,而语用意图在后期巩固,使得\u201c轨迹形状\u201d比任何单层快照更具信息量。此外,基于LDA的有害方向(逐层学习)在随机划分中保持稳定(成对余弦相似度>0.97),支持投影序列作为可复现的结构化信号。基于HPD,我们引入了\u201cHerald\u201d(通过激活层动力学进行有害编码识别)。这种轻量级输入审核器从跨层投影序列中提取七维特征记录(斜率、曲率、单调性、起始层及相关统计量),并使用288参数的MLP进行分类。Herald每层存储一个d维方向(对于32层、d=4096的模型为262 KB),训练时无需梯度计算,推理时仅增加2.6×10⁻⁶的预填充FLOPs。在八个提示有害性基准和四个模型家族中,Herald在OLMo2-7B上实现了89.3的平均F1分数,在对抗性越狱检测上超越了所有测试的防护模型(98.4 vs. 96.9 F1),并在每个骨干网络上比先前的基于潜在变量的方法高出2.3-4.1个F1点。逐实例轨迹提供了机器可读的审计记录,揭示了有害性\u201c何时\u201d以及\u201c如何\u201d出现,相比单层方法具有可解释性优势。

英文摘要

We identify \textbf{Harmfulness Propagation Dynamics (HPD)}: for harmful prompts, the projection of the last-token hidden state onto a learned harm direction rises monotonically with transformer depth, whereas benign prompts remain flat or oscillatory. This cross-layer signature reflects harmful intent as a \emph{progressively resolved} semantic property: surface form appears early, while pragmatic intent consolidates later, making the \emph{trajectory shape} more informative than any single-layer snapshot. Moreover, LDA-based harm directions, learned per layer, remain stable across random splits (pairwise cosine similarity $>0.97$), supporting the projection sequence as a reproducible structured signal. Building on HPD, we introduce \textbf{\herald{}} (\textbf{H}armful \textbf{E}ncoding \textbf{R}ecognition via \textbf{A}ctivation \textbf{L}ayer \textbf{D}ynamics). This lightweight input moderator extracts a seven-dimensional feature record, slope, curvature, monotonicity, onset layer, and related statistics from the cross-layer projection sequence and classifies it with a 288-parameter MLP. \herald{} stores one $d$-dimensional direction per layer ($262$\,KB for a 32-layer, $d{=}4096$ model), requires no gradient computation during training, and adds only $2.6{\times}10^{-6}$ prefill FLOPs at inference. Across eight prompt-harmfulness benchmarks and four model families, \herald{} achieves an average F1 of $89.3$ on OLMo2-7B, surpassing all tested guard models on adversarial jailbreak detection ($98.4$ vs.\ $96.9$ F1) and outperforming prior latent-based methods by $2.3$-$4.1$ F1 points on every backbone. Per-instance trajectories provide machine-readable audit records that reveal \emph{when} and \emph{how} harmfulness emerges, offering an interpretability advantage over single-layer approaches.

发表机构

  • İTÜ(伊斯坦布尔理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑