发表机构
Rutgers University; University of Washington(罗格斯大学; 华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过机制分析揭示MLLMs计数失败源于个体化与聚合双瓶颈,并提出ConvStack轻量架构以零初始化残差连接注入局部空间结构,仅微调计数任务即显著提升密集计数与空间理解性能。
AI 中文摘要
多模态大语言模型(MLLMs)在细粒度视觉计数任务上持续表现不佳,但其根本原因仍未被充分理解。在本工作中,我们对这一失败模式进行了机制性分析,识别出MLLMs全局注意力流程中固有的两个关键瓶颈。首先,我们揭示了一个源于图像分块处理的个体化瓶颈:由于视觉变换器(Vision Transformers)独立处理图像块,它们难以将跨边界的碎片化几何特征分组为不同的对象表征。其次,我们发现了后续计数聚合过程中的坍缩现象:随着数量增加,注意力压缩导致表征分离度迅速下降。识别并形式化这两个孪生瓶颈构成了我们的第一项主要贡献。为克服这些瓶颈,我们提出了ConvStack,一种轻量级架构,它直接在视觉令牌空间中操作,通过零初始化残差连接显式聚合并注入局部空间结构。通过显式解决个体化瓶颈,ConvStack为下游聚合提供了明确的几何证据。值得注意的是,仅通过在计数任务上进行微调,该模型在密集物体计数和更广泛的空间理解基准上取得了显著改进,且未损害通用视觉能力。
英文摘要
Multimodal Large Language Models (MLLMs) consistently struggle with fine-grained visual counting, yet the underlying causes remain poorly understood. In this work, we present a mechanistic analysis of this failure mode, identifying two critical bottlenecks inherent to the global attention pipeline of MLLMs. First, we reveal an individuation bottleneck stemming from image patchification: because Vision Transformers process patches independently, they struggle to group fragmented geometric features across boundaries into distinct object representations. Second, we identify a collapse in the subsequent counting aggregation process, where representation separation rapidly diminishes as numerosity increases due to attention compression. Identifying and formalizing these twin bottlenecks constitutes our first major contribution. To overcome them, we propose ConvStack, a lightweight architecture that operates directly in the visual token space to explicitly aggregate and inject local spatial structures via zero-initialized residual connections. By explicitly addressing the individuation bottleneck, ConvStack provides unambiguous geometric evidence for downstream aggregation. Remarkably, by fine-tuning exclusively on counting tasks, the model achieves substantial improvements in dense object counting and broader spatial understanding benchmarks, without compromising on general visual capabilities.