发表机构
Shenzhen University; The Hong Kong University of Science and Technology (Guangzhou); Hunan University of Science and Technology; Peng Cheng Laboratory; Fudan University(深圳大学; 香港科技大学(广州); 湖南科技大学; 鹏城实验室; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
UniCounting通过实例感知结构推理整合过完备提议,仅训练轻量关系头,在无图像查询多类别计数中实现更低误差,并揭示碎片化-合并权衡。
AI 中文摘要
视觉计数通常被表述为对单一指定目标的计数,模型接收特定图像的示例、文本查询或目标类别,并返回单一计数。我们转而研究固定词汇、无图像查询的多类别计数。每次运行时固定一个全局词汇表,仅给定RGB图像,模型预测完整的类别-计数向量,而无需被告知图像中出现哪些类别。我们提出UniCounting,将计数视为对过完备提议集合的实例感知结构推理。通用分割器会产生重复掩码、部分视图以及来自邻近实例的提议;语义分数可以为其命名,但无法确定哪些提议指向同一对象。冻结的SAM 2.1生成掩码,而冻结的DINOv2和OpenCLIP提供关系与类别特征。一个包含3,267个参数的类别共享关系头,基于实例掩码派生的监督信号预测同一实例的亲和度。随后通过稀疏图构建、代表性选择、标记和背景余量接纳,将每个被接纳的组件转换为一个计数,并附带可回放的组证据。仅训练关系头,无需计数或密度图目标。在COCO clean500上,UniCounting相比校准的OWLv2-All80,获得了更低的点估计向量ℓ1误差和更低的缺失类别假阳性质量,同时具有相当的微平均存在性F1分数。在匹配的解码器下,学习到的关系相对于掩码包含、掩码IoU、CLIP和DINO,同时减少了两种误差,并揭示了碎片化-合并的权衡。我们还报告了在OmniCount-sub、FSC-147和CARPK上的迁移诊断结果。
英文摘要
Visual counting is commonly formulated as counting a single specified target, with a model receiving an image-specific exemplar, text query, or target category and returning a single count. We instead study fixed-vocabulary image-query-free multi-category counting. A global vocabulary is fixed for each run, and, given only an RGB image, the model predicts a complete category--count vector without being told which categories appear. We present UniCounting, which casts counting as instance-aware structural inference over an over-complete proposal set. Generic segmenters produce duplicate masks, partial views, and proposals from neighboring instances; semantic scores can name them but cannot determine which denote the same object. Frozen SAM~2.1 generates masks, while frozen DINOv2 and OpenCLIP provide relation and category features. A 3,267-parameter category-shared relation head predicts same-instance affinities from instance-mask-derived supervision. Sparse graph construction, representative selection, labeling, and background-margin admission then convert each admitted component into one count with replayable group evidence. Only the relation head is trained, without count or density-map targets. On COCO clean500, UniCounting obtains lower point-estimate vector $\ell_1$ error and absent-class false mass than calibrated OWLv2-All80, with comparable micro presence F1. Under a matched decoder, the learned relation reduces both errors relative to mask containment, mask IoU, CLIP, and DINO, while revealing a fragmentation--merge trade-off. We also report transfer diagnostics on OmniCount-sub, FSC-147, and CARPK.