arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过提示随机Transformer实现无训练的通用逼近

Training-Free Universal Approximation by Prompting Random Transformers

Alexander Hsu, Rongjie Lai

arXiv 2608.09558首次发表:更新:

发表机构

Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究证明预训练Transformer可选,随机未训练的单层softmax注意力网络经软提示引导可实现通用逼近,其继承核回归的极小极大最优速率,还揭示了提示相关参数的权衡关系。

AI 中文摘要

提示Transformer的表达能力如何?回答该问题对区分Transformer模型中提示、架构与预训练的作用,以及确定任务特定行为是否必须存储在模型权重中、还是可在推理时通过提示诱导至关重要。我们从逼近理论角度证明,预训练是可选的:带有随机未训练权重的单层softmax注意力网络,在适当软提示引导下,可逼近紧流形上的任意Hölder函数。基于softmax注意力与核方法的关联,我们构造显式软提示(每个目标函数对应一个提示,与查询无关),作为将注意力logits匹配到高斯核指数的线性系统的解,在此构造下,冻结的Transformer可模拟经典Nadaraya-Watson核估计器。该构造仅需权重满足温和秩条件,我们证明高斯初始化下该条件几乎必然成立。受提示网络继承核回归的理论保证,得到依赖本征维数、达到极小极大最优速率的通用逼近定理。我们进一步量化提示的代价,揭示构造的软提示token的范数、提示长度与隐藏维数间存在权衡关系。数值实验验证了上述构造及预测的速率。

英文摘要

How expressive is prompting a transformer? Answering this question is important for separating the roles of prompting, architecture, and pretraining in transformer models, and for determining whether task-specific behavior must be stored in model weights or can instead be induced at inference time through the prompt. We show, in an approximation-theoretic sense, that pretraining is optional: a single-layer softmax attention network with random, untrained weights can approximate any Hölder function on a compact manifold when steered by an appropriate soft prompt. Guided by the connection between softmax attention and kernel methods, we construct explicit soft prompts (a prompt per target function, independent of the query) as solutions to linear systems matching attention logits to Gaussian kernel exponents, under which the frozen transformer emulates the classical Nadaraya-Watson kernel estimator. The construction requires only a mild rank condition on the weights, which we show holds almost surely under Gaussian initialization. The prompted network inherits the theoretical guarantees of kernel regression, leading to universal approximation theorems with minimax-optimal rates that depend on the intrinsic dimension. We further quantify the cost of prompting, exposing a tradeoff between the norm of the constructed soft prompt tokens, prompt length, and hidden dimension. Numerical experiments corroborate the constructions and predicted rates.

Comments31 pages, 5 figures. Comments welcome!

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑