分量加权质心搜索用于精确增量式BPE
Component-Weighted Centroid Search for Exact Incremental BPE
浏览论文内容
中文总结 AI 辅助
针对精确增量式BPE,提出分量加权质心搜索,将每次追加时间从$O(\log^2 t)$降至$O(\log t)$,保持语义与空间复杂度,并提供最坏情况保证。
中文摘要 AI 辅助
精确增量式BPE在每次追加字节后维护规范的词元化状态。Jiang和Gong(2026)的最新算法在最坏情况下以$O(\log^2 t)$时间完成此操作,其中$t$是最大规范词元长度。其质心搜索访问$O(\log t)$个分量,并且每个分量可能额外花费$O(\log t)$进行有序点定位。在Jiang和Gong的归一化/适当地合并阶段模型中,我们仅更改该局部搜索。每个区间按其选择的递归分量的大小加权,因此从大小$m$移动到大小$m'$的成本为$O(1+\log(m/m'))$。这些费用呈望远镜式累加,使得每次追加的时间为$O(\log t)$,在$n$字节流上总时间为$O(n\log t)$,同时保持相同的BPE语义和渐近空间复杂度。我们还构造了一个固定字母表上的归一化适当BPE族,其中计数平衡搜索在可达更新上使用$\Theta(\log^2 t)$次探测,而加权搜索使用$\Theta(\log t)$次。Rust实现与每个测试实例的预测探测次数匹配。然而,在普通词汇表上,查询度数较小,因此改进是最坏情况保证而非平均速度结果。
英文摘要
Exact incremental BPE maintains the canonical tokenization state after every appended byte. The recent algorithm of Jiang and Gong (2026) does this in $O(\log^2 t)$ worst-case time, where $t$ is the maximum canonical token length. Its centroid search visits $O(\log t)$ components and can pay another $O(\log t)$ for ordered point location at each one. Within Jiang and Gong's normalized/proper merge-stage model, we change only that local search. Each interval is weighted by the size of the recursive component it selects, so a move from size $m$ to size $m'$ costs $O(1+\log(m/m'))$. These charges telescope, giving $O(\log t)$ time per append and $O(n\log t)$ over an $n$-byte stream, with the same BPE semantics and asymptotic space. We also construct a normalized proper BPE family over a fixed alphabet where count-balanced search uses $Θ(\log^2 t)$ probes on a reachable update, while the weighted search uses $Θ(\log t)$. A Rust implementation matches the predicted probe counts on every tested instance. On ordinary vocabularies the queried degrees are small, however, and the improvement is a worst-case guarantee rather than an average-speed result.
发表机构
- Yale University(耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。