发表机构
Guangdong University of Technology; School of Computer Science, Peking University(广东工业大学; 北京大学计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉Transformer处理空间冗余令牌的问题,提出边界少数自由能自适应筛选(BMFA)方法,通过构建分层近似并递归细化块来改进,经多项实验验证,该方法能降低低估、缩小差距并保持准确率。
AI 中文摘要
视觉Transformer只有在粗略的令牌摘要保留指数注意力聚合所需的证据时,才能有效地处理空间冗余令牌。我们识别出一种边界少数低估失败情况,即空间上小的高响应区域贡献了主导的吉布斯质量,而对块均值几乎不可见。我们通过归一化对数平均指数自由能与均值汇总之间的差异来形式化这种失败,证明少数吉布斯质量在其空间支持和均值贡献消失时仍可保持非零,并刻画了有限阶矩校正的局限性。在此分析基础上,我们引入了边界少数自由能自适应筛选(BMFA),它构建了分层分段常数近似,并根据局部自由能的可计算下限增量递归地细化块。通过控制合成测试、COCO和LVIS诊断探针、闭环DeiT-Tiny评估以及ImageNet-1K实验建立了一致的证据链。BMFA在叶比率为5.794%时将平均合成低估从2.582降低到0.261,将COCO图像边缘平均差距从2.254降低到0.526,并在叶比率为55.861%时保持71.520%的ImageNet Top-1准确率。当前原型在全QK计算后评估选择质量;因此报告的叶比率表征的是表示粒度而非验证的稀疏内核加速。
英文摘要
Vision Transformers process spatially redundant tokens efficiently only when coarse token summaries preserve the evidence required by exponential attention aggregation. We identify a boundary-minority underestimation failure in which a spatially small, high-response region contributes dominant Gibbs mass while remaining nearly invisible to a block mean. We formalize the failure through the discrepancy between normalized log-mean-exp free energy and mean summarization, prove that minority Gibbs mass can remain non-vanishing as its spatial support and mean contribution vanish, and characterize the limitations of finite-order moment corrections. Building on the resulting analysis, we introduce Boundary-Minority Free-Energy Adaptive Screening (BMFA), which constructs a hierarchical piecewise-constant approximation and recursively refines blocks according to a computable lower-bound increment of local free energy. Controlled synthetic tests, COCO and LVIS diagnostic probes, closed-loop DeiT-Tiny evaluations, and ImageNet-1K experiments establish a consistent evidence chain. BMFA reduces the mean synthetic underestimate from 2.582 to 0.261 at a 5.794% leaf ratio, lowers the COCO image-edge mean gap from 2.254 to 0.526, and preserves 71.520% ImageNet Top-1 accuracy at a 55.861% leaf ratio. The current prototype evaluates selection quality after full QK computation; the reported leaf ratio therefore characterizes representation granularity rather than verified sparse-kernel speedup.
Comments14 pages. Experiments include synthetic analysis, COCO, LVIS, closed-loop DeiT-Tiny evaluation, and ImageNet-1K classification