下一词元泛函估计
Next-token functional estimation
浏览论文内容
中文总结 AI 辅助
本文提出留窗口估计器,用于在时间相依序列中估计下一词元泛函,以参数速率收敛,并优于留一法等基线。
中文摘要 AI 辅助
假设我们观测到一个长度为 $n+1$ 的随机变量序列的前 $n$ 个点,并希望估计未观测到的最后一点以及 $n$ 个观测训练点的经验测度的某个泛函。此类下一词元泛函包括下一词元是新词的概率(也称为惊奇概率)、下一词元与训练点之间最小距离的尾概率,以及在观测点上训练的分类器的测试误差。所有这些量经典上通过留一法估计,但该方法在时间相依性下是不一致的。我们提出了一种留窗口估计器,它在形成经验测度之前删除每个索引之后长度为 $\ au$ 的窗口,并在 $\ au = 1$ 时退化为留一法。在自然假设下,我们证明对于任何平稳的 $\eta$-混合过程,若其还允许 Marton 耦合,我们的估计器的误差以参数速率衰减。因此,我们的结果覆盖了一大类随机过程上的几个自然泛函。我们通过一个尖锐的极小极大下界来补充这些上界,该下界针对混合马尔可夫链上惊奇概率的估计。在马尔可夫链、滑动平均过程和自回归过程上的模拟表明,我们的估计器在许多留一法和加常数基线失败的场景中都能成功。
英文摘要
Suppose we observe the first $n$ points of a sequence of random variables having length $n+1$, and wish to estimate a functional of the unobserved final point and the empirical measure of the $n$ observed training points. Such next-token functionals include the probability that the next token is novel (also known as the surprise probability), the tail probability of the minimum distance between the next token and training points, and the test error of a classifier trained on the observed points. All of these quantities are classically estimated by the leave-one-out method, which is inconsistent under temporal dependence. We propose a leave-a-window-out estimator, which deletes a window of length $τ$ after each index before forming the empirical measure and reduces to leave-one-out at $τ= 1$. Under natural assumptions, we show that the error of our estimator decays at a parametric rate for any stationary $β$-mixing process that also admits a Marton coupling. Our results thus cover several natural functionals on a large class of stochastic processes. We complement these upper bounds with a sharp minimax lower bound for estimating the surprise probability on mixing Markov chains. Simulations on Markov chains, moving-average processes, and autoregressive processes show that our estimator succeeds in many scenarios where leave-one-out and add-constant baselines fail.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。