发表机构
Beijing University of Posts and Telecommunications; JIUTIAN Research; Nanyang Technological University; Chongqing University of Posts and Telecommunications(北京邮电大学; 中移九天; 南洋理工大学; 重庆邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出逐令牌残差比较(TRC)方法,通过分析残差流动态定位并抑制重复相关异常,在LVLMs、LLMs和LRMs中平均降低循环率57%,为缓解非受控重复及资源消耗攻击提供早期干预途径。
AI 中文摘要
非受控重复会延长大型语言模型(LLMs)的自回归生成过程,并可能引发资源消耗攻击。先前对重复生成的分析已在中间层和深层识别出强烈激活的特征。然而,非受控重复活动在这些层中变得显著之前是如何出现和发展的,仍未被充分理解。本文主要在大视觉语言模型(LVLMs)中研究这一问题,这些模型通过视觉和文本输入支持更丰富的非受控重复类型。我们提出逐令牌残差比较(Tokenwise Residual Comparison, TRC)方法,该方法通过生成过程中的残差动态识别并定位与重复相关的异常。TRC比较跨生成令牌的注意力机制和多层感知机对残差流的写入,以识别与重复相关的模式,然后在识别出的层中选择性抑制残差流中的坐标。实验表明,TRC能有效缓解非受控重复,平均将循环率降低57%。我们的分析进一步表明,重复语义在浅层出现并通过残差流传播,干扰正常表示。TRC也适用于大型语言模型(LLMs)和大型推理模型(LRMs),在这些模型中它一致地捕获类似的重复动态并实现有效缓解。我们的工作将重复生成的研究从其显著的内部表示扩展到更早的干预机会,为缓解资源消耗攻击提供见解。
英文摘要
Uncontrolled repetition can prolong autoregressive generation in large language models (LLMs) and enable resource consumption attacks. Prior analyses of repetitive generation have identified strongly activated features in intermediate and late layers. However, how uncontrolled repetition activity emerges and develops before becoming prominent in these layers remains insufficiently understood. In this paper, we investigate this question primarily in large vision-language models (LVLMs), which support a richer set of uncontrolled repetitions through both visual and textual inputs. We propose Tokenwise Residual Comparison (TRC), a method that identifies and localizes anomalies associated with repetition from residual dynamics during generation. TRC compares attention and multilayer perceptron writes to the residual stream across generated tokens to identify patterns associated with repetition. It then selectively suppresses coordinates in the residual stream at the identified layer. Experiments show that TRC effectively mitigates uncontrolled repetition, reducing loop rates by 57\% on average. Our analysis further shows that repetition semantics emerge in shallow layers and propagate through the residual stream, disrupting normal representations. TRC also generalizes to large language models (LLMs) and large reasoning models (LRMs), where it consistently captures analogous repetition dynamics and achieves effective mitigation. Our work broadens the study of repetitive generation from its prominent internal representations to earlier opportunities for intervention, providing insights for mitigating resource consumption attacks.