发表机构
University of Doha for Science and Technology(多哈科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉语言模型人群计数精度不足的问题,提出CrowdCue方法,通过将专家模型的整数计数以文本或视觉符号形式条件化输入VLM,显著提升计数精度,最佳MAE达62.65。
AI 中文摘要
生成式视觉语言模型(VLM)提供了一种计数范式,其中单个模型同时产生计数结果和对场景的自然语言描述,然而其原始计数精度仅处于亚百万参数专用回归器的水平。一个开放的问题是,来自预训练专家的辅助提示能否将其提升到实用范围,以及该提示应通过何种通道进行路由。我们在四个广泛使用的人群计数基准(ShanghaiTech A和B、UCF-QNRF、NWPU-Crowd)上评估了Qwen2.5-VL-7B。零样本提示很少产生可解析的计数,因此LoRA监督微调建立了基线,整体MAE为81.64。将P2PNet派生的密度热图作为辅助视觉信号进行条件化,在我们测试的所有编码方式中均失败,而对抗性交换协议表明模型读取了热图但适得其反地应用了它。我们提出CrowdCue,一个将同一专家已整合的整数计数作为离散符号提供给VLM的系列。文本通道变体达到MAE 72.04。视觉通道变体将整数渲染为打印数字并作为第二张图像提供,达到MAE 62.65,这是本文中最强的结果,且远超提供提示的专家单独的结果(在同一分割上为84.45)。在我们研究的晚期融合VLM中,约束因素不是通道,而是专家信号传递的抽象级别。
英文摘要
Generative vision-language models (VLMs) offer a counting paradigm in which one model produces both a count and a natural-language account of the scene, yet their raw counting accuracy sits in the range of sub-million-parameter specialist regressors. The open question is whether auxiliary guidance from a pretrained specialist can lift them into useful territory, and through which channel that guidance is best routed. We evaluate Qwen2.5-VL-7B on four widely used crowd counting benchmarks (ShanghaiTech A and B, UCF-QNRF, NWPU-Crowd). Zero-shot prompting rarely produces a parseable count, so LoRA supervised fine-tuning establishes the baseline at overall MAE 81.64. Conditioning on a P2PNet-derived density heatmap as an auxiliary visual signal fails in every encoding we tested, and an adversarial-swap protocol shows the model reads the heatmap but applies it counterproductively. We propose CrowdCue, a family that supplies the same specialist's already-integrated integer count to the VLM as a discrete symbol. The text-channel variant reaches MAE 72.04. The visual-channel variant, which renders the integer as printed digits and supplies it as a second image, reaches MAE 62.65, the strongest result in this paper and well ahead of the cue-supplying specialist alone (84.45 on the same split). In the late-fusion VLM we study, the binding constraint is not the channel but the abstraction level at which the specialist signal is delivered.