arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

说得越多,付得越多:LLM服务中供应商侧Token膨胀的黑盒审计

The More It Says, the More You Pay: A Black-Box Audit of Provider-Side Token Inflation in LLM Services

Leilei Chen, Lan Zhang, Chen Tang, Pengcheng Sun, Jiewei Lai, Yixiao Huang, Zhaopeng Zhang, Xinpeng Shen

arXiv 2609.20370首次发表:更新:

发表机构

University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对按Token付费的LLM服务中供应商隐蔽膨胀输出Token的攻击(PTIA),本文提出一种基于饱和现象的单探针黑盒审计方法,无需参考模型,检测率达85.1%,假阳性率低于2%。

AI 中文摘要

在按Token付费的LLM服务中,模型说得越多,用户支付得越多。不诚实的供应商可以隐蔽地操纵生成过程以膨胀输出Token,同时在很大程度上保持任务效用。我们将这种操纵定义为供应商侧Token膨胀攻击(PTIA),并在供应商控制管道的查询、提示、表示和模型四个层面实例化了五种代表性攻击。我们的实验表明,每种攻击都能将平均输出长度增加到干净基线的10.2倍以上,证明了PTIA在生成多个阶段的经济吸引力和可行性。然而,从黑盒响应中审计PTIA对用户来说很困难。我们的关键观察是PTIA饱和现象:初始攻击会急剧延长输出,但进一步强化或组合的效果则小得多。我们将这种饱和归因于停止行为:初始PTIA会急剧降低序列结束(EOS)Token的概率,而进一步干预仅会略微降低该概率。基于这一见解,我们设计了一种轻量级的单探针审计方法,该方法应用一种受控的延长干预。在PTIA下,探针诱导的额外Token远少于正常服务。该审计既不需要可信的本地参考模型,也不需要历史干净响应,并且其分别发出的原始请求和探针请求类似于普通流量,使得规避变得困难。在四个开放权重模型上,它实现了平均85.1%的检测率,假阳性率低于2%。在15个真实LLM API服务中,该审计标记了7个具有PTIA一致行为的服务。

英文摘要

In pay-per-token LLM services, the more a model says, the more users pay. Dishonest providers can covertly manipulate generation to inflate output tokens while largely preserving task utility. We define such manipulation as a Provider-Side Token Inflation Attack (PTIA) and instantiate five representative attacks at the query, prompt, representation, and model levels of the provider-controlled pipeline. Our experiments show that each attack increases mean output length to more than 10.2x the clean baseline, demonstrating PTIA's financial appeal and feasibility at multiple stages of generation. Yet auditing PTIA from black-box responses is difficult for users. Our key observation is PTIA saturation: an initial attack sharply lengthens output, but further strengthening or composition has much less effect. We trace this saturation to stopping behavior: an initial PTIA sharply lowers the end-of-sequence token probability, whereas further intervention lowers it only marginally. Building on this insight, we design a lightweight single-probe audit that applies a controlled lengthening intervention. Under PTIA, the probe induces far fewer additional tokens than under normal service. The audit requires neither a trusted local reference model nor historical clean responses, and its separately issued original and probed requests resemble ordinary traffic, making evasion difficult. Across four open-weight models, it achieves an average detection rate of 85.1% with false-positive rates below 2%. Across 15 real LLM API services, the audit flags 7 for PTIA-consistent behavior.

Comments22 pages, 10 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑