arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

令牌原生存储:以智能体的语言进行读写

Token-Native Storage: Read and Write in your Agent's Language

Kumar Shivendu

arXiv 2608.02376首次发表:更新:

发表机构

Qdrant(Qdrant)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出令牌原生存储思路,将文本以模型的BPE令牌ID保存,压缩比和读写速度均优于UTF-8等字节编码,呼吁AI实验室标准化令牌化器以推进该方案。

AI 中文摘要

搜索引擎和数据库引擎仍将文本存储为专为人类设计的UTF-8格式,但越来越多读写这些文本的系统(嵌入器、重排序器和语言模型智能体)使用令牌ID而非字符,因此每次访问都需要在两者之间进行转换。随着智能体成为存储文本的主要读写者,本文提出令牌原生存储的思路:将文本以模型自身的字节对编码(BPE)令牌ID形式保存,这种方式既更小又更快。将r50k令牌ID打包为uint16时,在英文文本上不压缩的情况下已比UTF-8小2.25倍,使用熵编码器则可达3.30倍。在6种令牌化器和3种语料库(英文、代码、印地语)上,压缩令牌ID的效果优于或媲美所有字节编解码器,甚至包括语料库训练的zstd字典。两项发现进一步支持该思路:BPE按合并顺序而非频率对令牌编号,通过频率重排序,普通整数编解码器(streamvbyte)可恢复熵编码器的大部分压缩比,同时解码速度提升约7倍,这是一行代码的修改,我们呼吁AI实验室在发布词汇表时进行该修改。此外,由于模型读取的是令牌ID而非文本,令牌原生存储可直接将令牌ID交付给模型,无需在每次读取时重新令牌化,速度提升约10至600倍。唯一的障碍是共享令牌ID需要通用令牌化器,目前不同模型家族间尚未完全实现,因此我们呼吁标准化:发布并共享统一词汇表,如同ASCII和UTF-8对文本的标准化。

英文摘要

Search and database engines still store text as UTF-8, a format built for humans. But the systems that increasingly read and write that text (embedders, rerankers, and language-model agents) work with token IDs, not characters, so every access pays to translate between the two. As agents become the primary readers and writers of stored text, we argue for token-native storage: keep the text as the model's own byte-pair-encoding (BPE) token IDs. Packing r50k IDs as uint16 already beats UTF-8 by 2.25x on English with no compression, and an entropy coder on top reaches 3.30x. Across six tokenizers and three corpora (English, code, Hindi), compressing token IDs matches or beats every byte codec, even a corpus-trained zstd dictionary. Two findings sharpen the case. BPE numbers tokens by merge order instead of frequency, and re-ranking by frequency lets a plain integer codec (streamvbyte) recover most of the entropy coder's ratio while decoding ~7x faster, a near-free change to how AI labs publish vocabularies. And because a model reads token IDs, not text, a token-native store hands over the IDs directly instead of re-tokenizing on every read. The only requirement is that reader and writer share a tokenizer, and different model families often use different ones today, so we argue for standardization: a published, shared vocabulary, the way ASCII and UTF-8 standardized text.

Comments12 pages, 6 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑