arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19510cs.LGcs.DSstat.MEstat.ML

自回归模型中的总变差距离估计

Total Variation Distance Estimation in Autoregressive Models

  • University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

Eric Price, Kevin Tian, Zhiyang Xun, Yusong Zhu

AI总结:

研究自回归模型中两个分布的总变差距离估计问题。在样本、对数概率、噪声对数概率三种访问模型下,分别给出不同查询次数的估计方法,并通过实验验证算法的稳健性和实用性,即便KL散度无穷时也能估计。

AI中文摘要:

现代大语言模型部署在固定权重之上使用多种实现选择和推理优化(如批处理、自定义内核和量化),导致两个运行“相同模型”的引擎可能产生显著不同的分布。我们研究在三种访问模型下,估计两个长度为n的自回归分布之间的总变差(TV)距离至加性误差ε的问题。在样本访问下,使用$\widetilde{O}(n^2K/\varepsilon^2)$次查询;在对数概率访问下,使用$O(n/\varepsilon^2)$次查询且该结果是紧的;在噪声对数概率访问下,若概率值有相对误差σ,则使用$\widetilde{O}((n + n^2\sigma^2)/\varepsilon^2)$次查询。通过对算法的实证评估补充理论结果,实验突出了估计总变差距离的稳健性和实用性,即使在KL散度无穷时也可估计。代码可通过给定链接获取。

英文摘要:

Modern LLM deployments use a number of implementation choices and inference optimizations (e.g., batching, custom kernels, and quantization) on top of fixed weights, so two engines serving "the same model" can produce meaningfully different distributions. We study the problem of estimating the total variation (TV) distance between two length-$n$ autoregressive distributions to additive error $\varepsilon$, under three access models. (1) Under sample access, we use $\widetilde{O}(n^2 K/\varepsilon^2)$ queries, where $K$ is the maximum support of the next-token distribution. This improves upon the $\widetilde{O}(n^3 m/\varepsilon^5)$-query estimator of Meel et al. (2025), where $m \geq K$ is the total size of the token alphabet. (2) Under logit access, we use $O(n/\varepsilon^2)$ queries, and this is tight. (3) Under noisy logit access, we smoothly interpolate between the above two guarantees: if probability values are given to relative error $σ$, we use $\widetilde{O}((n+n^2σ^2)/\varepsilon^2)$ queries. We complement our theoretical results with an empirical evaluation of our algorithms, for example measuring the distance between SGLang and vLLM serving identical weights. Our experiments highlight the robustness and practicality of estimating the total variation distance, which remains estimable where the KL divergence is infinite. Our code is available at https://github.com/XunZhiyang/llm-tv-estimation.

补充信息

↑