arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨模态物体计数:分类体系、基准、应用及开放挑战

Object Counting Across Modalities: Taxonomies, Benchmarks, Applications, and Open Challenges

Joana Konadu Owusu, Shivanand Venkanna Sheshappanavar

arXiv 2608.23845首次发表:更新:

发表机构

University of Wyoming; Geometric Intelligence Research Lab.(怀俄明大学; 几何智能研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述针对跨模态物体计数领域,提出五轴分类体系,梳理应用领域挑战,给出路线图,强调需构建稳健评估基础设施区分开放世界泛化与基准优化。

AI 中文摘要

物体计数方法已从特定类别的密度回归快速转向由开放词汇基础模型支撑的计数模型,目前可基于各类视觉和文本提示枚举实例。尽管这一转变标志着重大概念进步,但本综述认为,通用通用性的相关主张已超出评估基础设施的支撑能力。多数进展指标依赖少数饱和基准,模型可利用这些基准的统计规律;新推出的诊断数据集则揭示了模型在语义 grounding、时间身份及带遮挡的空间推理方面的系统性缺陷。为解决这些缺陷,本研究提出五轴分类体系(模态、机制、提示、监督级别、泛化设置),并据此对显微学、遥感、人群计数、农业等应用领域的文献进行审计,将普遍存在的挑战明确为六大结构性矛盾。基于此,本研究提出组合场景理解、主动计数智能体及统一多模态评估协议的路线图,核心要务是构建稳健的评估基础设施,以区分开放世界泛化与特定基准优化,而非简单的增量式工程改进。

英文摘要

Object-counting methods have rapidly shifted from class-specific density regression to open-vocabulary, foundation-model-backed counters. These methods now enumerate instances from various visual and textual prompts. While this shift marks major conceptual progress, our survey argues that claims of universal generality have outpaced the evaluative infrastructure. Most progress metrics rely on a few saturated benchmarks that models exploit for statistical regularities. Newly introduced diagnostic datasets reveal systematic failures in semantic grounding, temporal identity, and spatial reasoning with occlusion. To address these failures, we introduce a five-axis taxonomy (modality, mechanism, prompting, supervision level, and generalization setting). We use this taxonomy to audit the literature across application domains, including microscopy, remote sensing, crowd counting, and agriculture. This formalizes prevailing challenges into six structural contradictions. From these, we propose a roadmap for compositional scene understanding, active counting agents, and unified multimodal evaluation protocols. The main imperative is to build a robust evaluation infrastructure to distinguish open-world generalization from benchmark-specific optimization, rather than simple incremental engineering.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑