arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TokTier:面向智能体大语言模型服务的精确有状态分词方案

TokTier: Exact Stateful CPU+GPU Tokenization for Agentic LLM Serving

Zhenyu Zhang, Zhichao Cao

arXiv 2607.29678首次发表:更新:

发表机构

Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出TokTier,一种面向智能体LLM服务的精确有状态分词方案,通过增量修复、GPU加速分词等技术,将首次令牌生成时间降低16%-34%,饱和请求速率提升至1821次/秒,性能远超现有方案。

AI 中文摘要

大语言模型服务系统会缓存提示词的键值状态,但大多数前端仍会在每次调用时重新对完整请求文本进行分词。这一成本主要由编码智能体承担,这类智能体在每次获取小工具结果后会重新提交长对话文本,且难以实现复用,因为即使是短的追加内容也会改变前序序列末尾附近的分词边界。在来自两个智能体生态系统的153951次调用中,中位数调用追加约1400个字符,仅有1.0%-3.6%的调用会启动或重建包含数百万字符上下文的会话。在94.1%的 fleet 提示词缓存命中率下,分词耗时占首次生成令牌时间的比例最高达64%。TokTier是一种有状态分词服务,其核心约定是:生成的令牌ID始终与请求文本的完整参考分词结果完全一致。对于会话续传,它会对追加内容周围的小窗口进行重新分词,并在每次请求时执行稳定边界检查后仅拼接符合要求的部分;若检查失败则扩大窗口或回退至完整分词。对于无可复用前缀的调用,它会将GPT系列模型的正则预分词分解为本地运行规则,并在GPU上执行精确预分词和BPE分词。采样影子验证器会对实时流量进行重新检查。针对17种分词器系列,差分测试覆盖了1.5×10^10次拆分检查、12.4TB真实文本语料库以及93000+次重放的智能体步骤,且未出现任何分歧。增量修复在处理10万至300万字符时耗时0.5-1.1毫秒,比HF分词器快达437倍,在处理100万字符时比最强的基于缓存的基准方案Gigatoken(完全预热)快2.1倍。GPU完整分词可在0.87毫秒内对100万至1000万字符的请求进行编码,比HF分词器快达491倍,比已发表最快的CPU方法快23.4倍。结合vLLM,在记录的突发流量下,中位数首次令牌生成时间降低16%-34%,P99降低23%。在50毫秒P99目标下,4个修复核心加1个GPU可维持1821次请求/秒,而16核无状态前端的饱和速率仅为40次请求/秒。

英文摘要

LLM serving caches prompt KV state, yet most front ends still re-tokenize the full request on every call. Coding agents pay most: sessions repeatedly submit a long transcript after a small append, which can shift token boundaries near the end of the prior sequence. Across 153,951 calls the median append is ~1.4K characters; only 1.0-3.6% of calls start or rebuild a session, yet those carrymulti-million-character contexts. Fleet prompt-cache hit rate is 94.1%, and as it approaches 0.99, tokenization grows from 10% to 64% of time to first token (TTFT) in component measurements. TokTier is a stateful CPU+GPU tokenization service for this two-mode workload, under one contract: emitted token IDs are always identical to full reference tokenization. For session continuations it re-tokenizes a small window around the append and splices only when a per-request check finds a stable pre-tokenization boundary; failed checks widen the window or fall back to full reference tokenization. For calls without a reusable prefix it runs exact GPT-family regex pre-tokenization and BPE on a GPU. A sampled shadow verifier re-checks live traffic. Across 17 production tokenizer families, differential campaigns cover 1.5x10^10 split checks, a 12.4 TB real-text corpus, and 93,000+ replayed agent steps, with zero divergence. Incremental repair takes 0.5-1.1 ms from 100K to 3M characters, up to 437x faster than HF tokenization and 2.1x faster at 1M characters than the strongest cache-based baseline (Gigatoken) fully prewarmed. GPU tokenization encodes a 1M-character request in 0.87 ms, up to 491x below HF and 23.4x below the fastest published CPU method on the same protocol. With vLLM, median TTFT drops 16-34% and P99 TTFT 23% under recorded bursts. Under a 50 ms P99 objective, a four-core repair pool plus one GPU sustains 1,821 requests/s, where a 16-core stateless front end saturates at 40 requests/s.

Comments26 pages. Code: https://github.com/asu-idi/toktier

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑