arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

空间感知类无关目标计数

Spatially-Aware Class-Agnostic Object Counting

Robert Wijaya, Md. Tanvir Hossain, Amanda Kau, Ngai-Man Cheung

arXiv 2607.16826首次发表:更新:

发表机构

Singapore Univeristy of Technology and Design; The University of Canterbury; Heineken GenAI Lab(新加坡科技设计大学; 坎特伯雷大学; 喜力通用人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究广义目标计数问题,提出UpCount方法,通过提取多层特征、重组多尺度金字塔并利用相关技术增强空间敏感性,在FSC-147和CARPK数据集上取得较好计数效果,有效解决复杂目标计数难题。

AI 中文摘要

广义目标计数旨在从单张图像估计任意目标类别的实例数量,但由于空间建模有限,许多最新方法在结构复杂的目标上存在困难。我们提出了UpCount,这是一种旨在更好地保留空间结构的类无关计数器。UpCount通过从ViT-B/16编码器中提取多层特征并将它们重新组装成一个经过细化的多尺度金字塔来增强视觉表示,该金字塔使用密集预测变压器和FeatUp在空间上进行细化,从而产生具有改进的结构和空间敏感性的特征;然后,一个提议-验证计数头识别重复模式并生成用于最终计数的密度图。在FSC-147上,UpCount在测试集上实现了12.39的平均绝对误差(MAE)和100.89的均方根误差(RMSE),并且它有效地转移到了CARPK上的车辆计数(6.27 MAE,8.79 RMSE)。代码:此https URL

英文摘要

Generalised object counting aims to estimate the number of instances of an arbitrary object category from a single image, but many recent methods can struggle on structurally complex objects due to limited spatial modelling. We present \textit{UpCount}, a class-agnostic counter designed to better preserve spatial structure. UpCount strengthens the visual representation by extracting multi-layer features from a ViT-B/16 encoder and reassembling them into a refined multi-scale pyramid that is spatially refined using Dense Prediction Transformers and FeatUp, yielding features with improved structural and spatial sensitivity; a proposal--verification counting head then identifies repeated patterns and produces a density map for the final count. On FSC-147, UpCount achieves 12.39 MAE and 100.89 RMSE on the test set, and it transfers effectively to vehicle counting on CARPK (6.27 MAE, 8.79 RMSE). Code: https://github.com/r28112072-rgb/upcount

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑