Count Anything
Count Anything
- Tsinghua University(清华大学)
- China University of Geosciences, Wuhan(武汉地质大学)
- State Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室)
- National Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心)
- Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出跨域文本引导的目标计数模型Count Anything,通过双粒度实例枚举和互补计数融合,在统一基准CLOC上实现多域泛化。
AI中文摘要:
尽管通用视觉模型取得了快速进展,目标计数仍然分散在特定领域的数据集和任务公式中。现有的计数模型通常针对人群、车辆、细胞、农作物或遥感目标等场景定制,因此难以跨类别、视觉域、目标尺度和密度分布进行泛化。在本文中,我们研究了跨域的文本引导目标计数,其中模型以图像和自然语言查询为输入,并返回一组基于实例的目标点,其基数给出计数。这种公式将类别条件计数与可解释的空间定位统一起来。为了支持这一设置,我们构建了CLOC,一个跨域大规模目标计数数据集,将多样化的公共数据源重组为统一的基准。CLOC涵盖六个视觉域:通用场景、遥感、组织病理学、细胞显微镜、农业和微生物学,包含约22万张图像、619个类别和1500万个目标实例。基于CLOC,我们提出了Count Anything,一个用于文本引导目标计数的通用模型。与主导计数模型的密度图方法不同,Count Anything采用离散实例点并执行双粒度实例枚举。区域级稀疏计数器为大而稀疏的目标提供目标级锚点,而像素级密集计数器通过密集点预测处理小、拥挤和弱边界目标。点中心监督策略能够从异构标注中学习,互补计数融合以无参数方式结合两个计数器。大量实验表明,Count Anything实现了强准确性和多域泛化,优于现有的开放世界计数方法。代码可在:https://github.com/Mengqi-Lei/count-anything 获取。
英文摘要:
Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing counting models are often tailored to scenarios such as crowds, vehicles, cells, crops, or remote-sensing objects, and thus struggle to generalize across categories, visual domains, object scales, and density distributions. In this paper, we study text-guided object counting across domains, where a model takes an image and a natural-language query as input and returns an instance-grounded set of target points whose cardinality gives the count. This formulation unifies category-conditioned counting with interpretable spatial localization. To support this setting, we construct CLOC, a Cross-domain Large-scale Object Counting dataset that reorganizes diverse public data sources into a unified benchmark. CLOC covers six visual domains: General Scene, Remote Sensing, Histopathology, Cellular Microscopy, Agriculture, and Microbiology, with about 220K images, 619 categories, and 15M object instances. Based on CLOC, we propose Count Anything, a generalist model for text-guided object counting. Unlike density-map-based methods, which dominate counting models, Count Anything adopts discrete instance points and performs dual-granularity instance enumeration. A Region-level Sparse Counter provides object-level anchors for large and sparse targets, while a Pixel-level Dense Counter handles small, crowded, and weakly bounded targets via dense point prediction. A point-centric supervision strategy enables learning from heterogeneous annotations, and Complementary Count Fusion combines both counters in a parameter-free manner. Extensive experiments show that Count Anything achieves strong accuracy and multi-domain generalization, outperforming existing open-world counting methods. Code is available at: https://github.com/Mengqi-Lei/count-anything.