arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

函数空间Transformer与自适应锚点

Function-Space Transformer with Adaptive Anchors

Guorui Sang, Pedram Rooshenas

arXiv 2609.38348首次发表:更新:

发表机构

University of Illinois Chicago(伊利诺伊大学芝加哥分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出函数空间Transformer(FST),通过预测锚点位置并递归细化特征,实现空间自适应连续表示,在PDE预测和图像分类中优于基线模型。

AI 中文摘要

许多形式的数据,包括物理场、几何形状和视觉信号,自然由连续域上的函数描述,但通过离散样本观测。在固定均匀网格上表示这些函数,需要在解析局部变化和增加跨域计算之间进行权衡。神经算子通过学习函数之间的映射来解决这种不匹配,而潜在注意力架构则为采样观测提供灵活处理。我们引入函数空间Transformer(FST),一种通过空间自适应连续潜在表示从函数中学习的框架。FST在锚点处存储特征,锚点的位置由输入观测预测,并通过函数空间交互递归细化这些锚点特征。这使得表示能够适应每个输入的空间组织,而非继承观测网格的空间组织,同时支持空间解析和有限维输出。在使用PDEBench Burgers和Darcy流的PDE解预测中,FST显著优于Perceiver IO基线(其潜在表示缺乏显式空间组织),并与傅里叶神经算子高度竞争。在ImageNet-1K上,FST比Vision Transformer基线实现更高的分类准确率,且在这些比较中参数更少。消融研究进一步支持函数空间更新和递归细化的优势。这些结果共同凸显了自适应连续表示在科学预测和视觉识别中的潜力。

英文摘要

Many forms of data, including physical fields, geometric shapes, and visual signals, are naturally described by functions over continuous domains but are observed through discrete samples. Representing these functions on fixed uniform grids imposes a trade-off between resolving localized variation and increasing computation across the domain. Neural operators address this mismatch by learning mappings between functions, while latent-attention architectures provide flexible processing of sampled observations. We introduce the Function-Space Transformer (FST), a framework for learning from functions through a spatially adaptive continuous latent representation. FST stores features at anchors whose locations are predicted from the input observations and recursively refines these anchor features through function-space interactions. This allows the representation to adapt its spatial organization to each input rather than inherit that of the observation grid, while supporting both spatially resolved and finite-dimensional outputs. On PDE solution prediction using PDEBench Burgers and Darcy flow, FST substantially outperforms the Perceiver IO baseline, whose latent representation lacks explicit spatial organization, and is highly competitive with the Fourier Neural Operator. On ImageNet-1K, FST achieves higher classification accuracy than the Vision Transformer baseline, with fewer parameters across these comparisons. Ablations further support the benefits of function-space updates and recursive refinement. Together, these results highlight the potential of adaptive continuous representations for both scientific prediction and visual recognition.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑