arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

非参数上下文学习在增长几何复杂度下的研究:Transformer的极小极大最优性与局部几何自适应性

Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers

Jaehee Seo, Jisu Kim

arXiv 2609.31458首次发表:更新:

发表机构

Seoul National University(首尔国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对几何异质数据,提出结构感知的两阶段softmax Transformer,实现非参数上下文学习的极小极大最优性与局部几何自适应性。

AI 中文摘要

Transformer已成为上下文学习(ICL)的核心架构,尤其在大语言模型中展现出最先进的性能。这一成功促使我们理解Transformer如何利用几何异质数据中与任务相关的结构。然而,现有的非参数ICL理论主要集中于欧几里得域或单流形模型。为弥补这一空白,我们研究了未知局部几何下的预测问题,该问题由样本量依赖的流形混合模型建模,这些流形具有异质的维度、光滑性和采样质量。在局部分离和小扰动条件下,我们建立了捕捉各分量总体难度的极小极大下界,并构造了一个具有匹配上界的Oracle切向局部多项式估计器。该估计器与一个结构感知的两阶段softmax Transformer相关联,该Transformer配备了几何预处理器和图表式简化局部多项式求解器。该Transformer在极小极大速率下实现了可忽略的近似误差,且深度为对数级、大小为多项式级。最后,我们推导了该类上近经验风险最小化器的上下文泛化界。综合这些结果,我们确定了所得预测器利用局部几何并达到总体极小极大速率的条件。

英文摘要

Transformers have become a central architecture for in-context learning (ICL), particularly through their state-of-the-art performance in large language models. This success motivates understanding how transformers exploit task-relevant structure in geometrically heterogeneous data. However, existing nonparametric ICL theory has largely focused on Euclidean domains or single-manifold models. To address this gap, we study the prediction problem under unknown local geometry, modeled by sample size-dependent mixtures of manifolds with heterogeneous dimensions, smoothness, and sampling masses. Under local separation and small-perturbation conditions, we establish a minimax lower bound capturing the aggregate difficulty of the components and construct an oracle tangent local-polynomial estimator with a matching upper bound. This estimator is connected to a structure-informed, two-stage softmax transformer with a geometric preconditioner and chartwise reduced local-polynomial solvers. The transformer achieves negligible approximation error relative to the minimax rate with logarithmic depth and polynomial size. Finally, we derive an in-context generalization bound for near empirical risk minimizers over this class. Together, these results identify conditions under which the resulting predictor exploits local geometry and attains the aggregate minimax rate.

Comments63 pages, 2 figures. Accepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑