FragToken:通过非规范令牌生成放大LLM推理成本
FragToken: Amplifying LLM Inference Costs through Noncanonical Token Generation
- University of Electronic Science and Technology of China(电子科技大学)
- City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对LLM推理成本攻击,提出FragToken训练时框架,利用非规范令牌序列增加解码步数,在四个模型上实现1.99-2.46的令牌膨胀比,同时保持模型效用。
AI中文摘要:
随着大语言模型(LLM)推理变得越来越昂贵,资源消耗攻击对模型提供商构成了日益严重的威胁。现有攻击通常通过诱导攻击者控制或触发的请求产生异常长或重复的输出,从而放大成本,这使得它们容易被检测,并且在良性流量占主导时限制了其部署范围内的广泛影响。在这项工作中,我们揭示了一个先前被忽视的、源于令牌序列到解码文本的多对一映射的令牌级攻击面。尽管标准LLM主要生成由其分词器诱导的规范令牌序列,但相同的文本也可以用更长的非规范序列来表示。这种表示灵活性为资源消耗攻击开辟了一条新途径:攻击者可以训练模型偏向此类序列,从而在不按比例增加可见响应长度的情况下,系统地增加自回归解码步骤的数量。然而,我们通过实验发现,直接最大化令牌碎片化会显著降低模型效用,产生明显的答案质量失败,从而削弱攻击的隐蔽性。为了应对这一挑战,我们提出了FragToken,一个训练时框架,结合了源模型自蒸馏、容量感知过滤与预算分配以及BPE对齐合并,在普通提示下诱导碎片化生成,同时基本保持模型效用。我们在三个基准上对四个LLM评估了FragToken。在这四个模型中,FragToken实现了三个基准的平均令牌膨胀比(TIR)在1.99到2.46之间,同时仅导致模型效用的轻微下降。我们的工作揭示了一种隐蔽的LLM供应链威胁,它无需大量攻击请求即可增加推理成本,同时基本保持效用。
英文摘要:
As large language model (LLM) inference becomes increasingly expensive, resource-consumption attacks pose a growing threat to model providers. Existing attacks typically amplify cost by inducing abnormally long or repetitive outputs on attacker-controlled or triggered requests, making them easier to detect and limiting their deployment-wide impact when benign traffic dominates. In this work, we uncover a previously overlooked token-level attack surface arising from the many-to-one mapping from token sequences to decoded text. Although standard LLMs predominantly generate the canonical token sequences induced by their tokenizers, the same text can also be represented by substantially longer non-canonical sequences. This representational flexibility exposes a new avenue for resource-consumption attacks: an attacker can train the model to favor such sequences, systematically increasing the number of autoregressive decoding steps without a proportional increase in visible response length. However, we empirically find that directly maximizing token fragmentation substantially degrades model utility, producing conspicuous answer-quality failures that undermine attack stealthiness. To address this challenge, we propose FragToken, a training-time framework that combines source-model self-distillation, capacity-aware filtering and budgeting, and BPE-Aligned Merging to induce fragmented generation under ordinary prompts while largely preserving model utility. We evaluate FragToken on four LLMs across three benchmarks. Across the four models, FragToken achieves a three-benchmark average token inflation ratio (TIR) ranging from 1.99 to 2.46, while causing only minor degradation in model utility. Our work reveals a covert LLM supply-chain threat that increases inference cost without requiring large volumes of attack requests while largely preserving utility.