arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过上下文学习的自适应均值估计:梯度流分析

Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis

Martin Eppert, Krishna Balasubramanian, Subhro Ghosh, Jason Klusowski, Yan Shuo Tan

arXiv 2610.07804首次发表:更新:

发表机构

National University of Singapore; University of California, Davis; Princeton University(新加坡国立大学; 加州大学戴维斯分校; 普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过梯度流分析,证明先验拟合网络在位置估计中通过上下文学习实现统计自适应性,在多种数据分布下达到接近最优的估计速率。

AI 中文摘要

先验拟合网络(PFNs),如TabPFN,目前在预测和估计任务上已能与成熟统计程序相媲美。一个自然的解释是PFNs具有统计自适应性,即对于一组异构模型,它们表现得几乎与针对真实数据生成模型量身定制的方法一样好,而无需被告知数据来自哪个模型。我们在一个受控的位置估计问题中研究这种自适应性是如何被学习到的。每个任务都是一个未标记的样本,其所属分布族是隐藏的:高斯数据需要取平均,误差阶为$n^{-1}$,而均匀数据最好通过其极值来估计,速率为更快的$n^{-2}$。我们还提供了一个对称高斯混合的例子,对于该混合,可以达到$\n^2_n/n$的速率。在标量输入上,softmax注意力计算经验累积生成函数的导数。因此,一个单一原语既提供了区分这些分布族的特征,又形成了在样本均值和中程值之间插值的估计器。我们通过softmax专家混合或门控线性单元(GLU)来组合注意力专家,并分析分段梯度流。使用$\tilde{\Omega}(n^{1+\epsilon})$个预训练任务,学习到的估计器在高斯任务上是渐近有效的,在均匀任务上达到最小最大速率的$n^{\epsilon}$倍以内,并在方差收缩机制下对混合分布达到阶最优。这些保证扩展到新的位置和更长的上下文。风险分解将专家误差、路由误差和归一化误差分开,这阐明了架构上的对比。Softmax门控强制执行归一化和精确的平移等变性,而GLU必须学习这些:其动态分为快速偏差消除和随后的慢速专家选择。端到端实验恢复了预测的特化。

英文摘要

Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks. A natural explanation is that PFNs have the property of statistical adaptivity, that is, they perform nearly as well as a method tailored to the true data-generating model for a heterogeneous set of models, while not being told which model the data comes from. We study how such adaptivity is learned in a controlled location-estimation problem. Each task is an unlabeled sample whose family is hidden: Gaussian data call for averaging, with error of order $n^{-1}$, whereas uniform data are best estimated from their extremes, at the faster rate $n^{-2}$. We also provide the example of a symmetric Gaussian mixture, for which a rate of $σ^2_n/n$ can be attained. On scalar inputs, softmax attention computes the derivative of the empirical cumulant-generating function. A single primitive therefore both supplies features that distinguish the families and forms estimators interpolating between the sample mean and the mid-range. We combine attention experts through either a softmax mixture of experts or a gated linear unit (GLU), and analyze stagewise gradient flow. With $\widetildeΩ(n^{1+ε})$ pretraining tasks, the learned estimator is asymptotically efficient on Gaussian tasks, within a factor $n^ε$ of the minimax rate on uniform tasks, and order-optimal on mixtures in a shrinking-variance regime. These guarantees extend to new locations and longer contexts. A risk decomposition separates expert error, routing error and normalization error, which clarifies the architectural contrast. Softmax gating enforces normalization and exact translation equivariance, whereas the GLU must learn it: its dynamics separate into fast bias removal followed by slow expert selection. End-to-end experiments recover the predicted specialization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑