arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越实例的计数:面向组-个体对象计数的基准

Counting Beyond Instances: A Benchmark for Group-Individual Object Counting

Rui Wang, Junyi Huang, Jiahui Li, Qiao Yu, Yixue Hao, Long Hu, Baoru Huang

arXiv 2609.04716首次发表:更新:

发表机构

School of Computer Science and Technology, Huazhong University of Science and Technology; Shanghai Artificial Intelligence Laboratory; University of Liverpool(华中科技大学计算机科学与技术学院; 上海人工智能实验室; 利物浦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有计数范式忽视语义单元的局限,提出组-个体对象计数新设置,构建含对应标注的BunchCount基准,提出计数单元引导的关系计数框架,实现组级计数性能提升并保留个体级计数能力。

AI 中文摘要

视觉计数通常在实例层面进行,旨在估计图像中查询类别的对象数量。然而,现实世界中的计数往往涉及由多个实例组成的更高层次语义单元,例如一串葡萄、一叠盘子或一双鞋。这暴露了现有计数范式的一个关键局限:现有范式主要关注计数对象,却在很大程度上忽视了计数的语义单元。我们提出了Group-Individual Object Counting(GIC,组-个体对象计数)这一新设置,要求模型在统一框架内同时对个体对象和语义组进行计数。为支持这一新任务,我们构建了BunchCount这一现实世界基准,包含1330张图像、89254个个体标注和11065个组标注。BunchCount提供了同一图像内配对的个体-组标注,并明确记录了每个组与其组成个体之间的包含关系。在BunchCount上的实验表明,当前先进的计数模型在个体实例计数上表现良好,但无法更准确地对语义组进行计数。为缓解语义粒度冲突,我们提出了计数单元引导的关系计数框架,该框架利用组-个体包含关系在训练期间正则化跨粒度表示。我们的方法大幅提升了组级计数性能,同时更好地保留了个体级计数能力,为超越实例的计数建立了强大基线。

英文摘要

Visual counting is commonly formulated at the instance level, aiming to estimate how many objects of a queried category appear in an image. However, real-world counting often involves higher-level semantic units formed by multiple instances, such as a bunch of grapes, a stack of plates, or a pair of shoes. This exposes a key limitation of existing counting formulations, which mainly focus on what to count, while largely overlooking at which semantic unit to count. We introduce Group-Individual Object Counting (GIC), a new setting that requires models to count both individual objects and semantic groups within a unified framework. To support this new task, we present BunchCount, a real-world benchmark with 1,330 images, 89,254 individual annotations, and 11,065 group annotations. BunchCount provides paired individual-group annotations within the same image and explicitly records containment relations between each group and its constituent individuals. Experiments on BunchCount show that current advanced counting models perform well on individual instances but fail to count semantic groups more accurately. To mitigate semantic granularity conflict, we propose a counting-unit guided relational counting framework, which exploits group-individual containment relations to regularize cross-granularity representations during training. Our method substantially improves group-level counting while better preserving individual-level counting ability, establishing a strong baseline for counting beyond instances.

Comments14 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑