发表机构
University of Delaware; Iowa State University(特拉华大学; 爱荷华州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过可观测性、可实现性和实现三个条件,精确刻画了固定连续前缀在冻结注意力下替代低秩适配器的能力,并给出了误差下限、最优解形式及数值稳定性结果。
AI 中文摘要
一个固定的连续前缀能否在注意力头保持冻结的情况下替代给定的低秩适配器?在本研究中,我们表明答案取决于适配器的目标,并通过三个条件来刻画。首先,可观测性:在单次因果读取中,每个独立的键-值前缀仅通过查询、注意力划分和值分子来感知内容,因此,如果目标在两个具有相同摘要的输入上产生不同结果,则在任何前缀长度下都会存在误差下限;范数上限将此下限扩展到几乎相同的摘要。其次,可实现性:在共同查询下,任何前缀恰好简化为两个聚合变量,范数受限的最优解是一个可达的二阶锥规划,在固定输出投影后也是如此;它将两个等范数的秩一值更新置于可编译性的相对两侧。第三,实现:在仿射查询暴露下,$2r$ 个有符号槽近似一个秩-$r$ 的值更新,但其值增长为 $O(\epsilon^{-3/2})$,并且该构造在 float64 中通过所有 400 次容差检查,但在 bfloat16 中仅通过 38 次。第一层 GPT-2 读取头在固定 token 和位置下满足共同查询条件,而无需限制激活;在三个这样的头上,受限最优解留下 18.4% 至 74.2% 的投影适配器效应未编译,且值-查询排序依赖于头。所有声明均涉及单个头部的局部近似,而非整个网络的等价性。
英文摘要
Can a fixed continuous prefix replace a given low-rank adapter while the attention head stays frozen? In this research, we show that the answer depends on the adapter's target through three conditions. First, observability: at one causal readout, every independent key--value prefix sees the content only through the query, attention partition, and value numerator, so a target that differs on two inputs with equal summaries incurs an error floor at every prefix length; norm caps extend this floor to nearly equal summaries. Second, realizability: at a common query, any prefix reduces exactly to two aggregate variables, and the norm-capped optimum is an attained second-order-cone program, also after a fixed output projection; it places two equal-norm rank-one value updates on opposite sides of compilability. Third, implementation: under affine query exposure, $2r$ signed slots approximate a rank-$r$ value update, but their values grow as $O(ε^{-3/2})$, and the construction passes all 400 tolerance checks in float64 yet only 38 in bfloat16. A first-layer GPT-2 readout with fixed token and position meets the common-query condition without clamping activations; at three such heads, the capped optimum leaves 18.4\% to 74.2\% of the projected adapter effect uncompiled, with a head-dependent value--query ordering. All claims concern local approximation at one head, not whole-network equivalence.
Comments26 pages, 4 figures