AI 中文总结
本研究通过受控实验对比不同视觉定位解码器,发现仅注意力的4块解码器(A4)性能与含FFN的4块解码器相当,8块仅注意力解码器(A8)可弥补A4的微小差距且更优,同时A4能减少参数量与延迟。
AI 中文摘要
当预训练视觉-语言模型(VLM)已编码图像与语言上下文后,视觉定位解码器中的前馈网络(FFN)是否会增加必要计算?我们在冻结VLM特征上对比了仅含注意力的4块解码器(A4)、匹配的含注意力加FFN的4块解码器(S4)以及8块仅注意力的参数控制模型(A8)。在RefCOCOg和Ref-Adv-s数据集上,A4的表现与S4相当或略优;FineCops-Ref数据集显示,A4在IoU@0.5指标上较S4存在0.52个百分点的小差距(95%置信区间[0.12,0.95],偏向S4),但A8可弥补该差距且超出S4 0.26个百分点。官方FineCops指标未显示差距呈单调增长。A4使可训练解码器参数减少44.4%,缓存解码器延迟降低10.1%,不过端到端延迟仍由主干网络主导。上述结果针对可训练的定位解码器,而非完整的仅注意力VLM。
英文摘要
Do feed-forward networks (FFNs) in visual grounding decoders add essential computation once a pretrained vision-language model has already encoded image and language context? We compare a four-block attention-only decoder (A4), a matched four-block attention-plus-FFN decoder (S4), and an eight-block attention-only parameter control (A8) over frozen VLM features. A4 matches or slightly exceeds S4 on RefCOCOg and Ref-Adv-s. FineCops-Ref reveals a small A4 deficit of 0.52 percentage points at IoU@0.5 (95% CI [0.12, 0.95] in favor of S4), but A8 recovers it and finishes 0.26 points above S4. Official FineCops levels do not show a monotonic increase in the gap. A4 reduces trainable decoder parameters by 44.4% and cached-decoder latency by 10.1%, although end-to-end latency remains backbone-dominated. These results concern the trainable grounding decoder, not a complete attention-only VLM.
Comments14 pages, 8 figures, 5 tables. Code and project page: https://github.com/TarunTomar122/attention-is-all-you-need-for-vlms