arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06236cs.CVcs.AI

拥挤场景中基于深度引导的视频目标计数

Depth-Guided Video Object Counting in Crowded Scenes

  • Harbin Institute of Technology (Weihai)(哈尔滨工业大学(威海))
  • City University of Hong Kong(香港城市大学)
  • University of Science and Technology of China(中国科学技术大学)
  • Harbin Institute of Technology Qingdao Research Institute(哈尔滨工业大学青岛研究院)

机构由 AI 辅助整理,请以论文原文为准。

Yuanjing Xu, Xinyan Liu, Weidong Chen, Zixuan Zou, Linhao Zhang, Zhuangzhe Meng, Antoni B. Chan, Weigang Zhang

AI总结:

该研究针对拥挤场景视频目标计数的局限性,提出DG-Det与后处理流水线、去重框架,发布新数据集,使MAE降62.01%且RMSE提升。

AI中文摘要:

我们的主要目标是推进拥挤场景下的视频目标计数,旨在基于给定的文本或视觉提示稳健计数目标类别的所有实例。现有方法依赖RGB信息,在拥挤和遮挡条件下判别能力受限。为解决该问题,我们提出了深度引导检测器(DG-Det)及通用后处理流水线。通过将深度线索与多尺度RGB-D交叉注意力、显式遮挡预测相结合,我们的方法提升了空间理解能力,在拥挤和遮挡场景中实现稳健检测。此外,我们引入统一去重框架以消除跨帧冗余计数。为推动未来研究,我们还发布了新的RGB-D视频目标计数数据集,包含深度信息及每序列多个目标类别。大量实验表明,与现有基线相比,我们的方法使平均绝对误差(MAE)降低62.01%,均方根误差(RMSE)也取得持续提升。我们在该httpsURL提供了源代码,在该httpsURL提供了数据集。

英文摘要:

Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts. Existing methods rely on RGB information, limiting their discriminative ability in crowded and occluded conditions. To address this, we propose a Depth-Guided Detector (DG-Det) along with a general post-processing pipeline. By integrating depth cues with multi-scale RGB-D cross-attention and explicit occlusion prediction, our method enhances spatial understanding and achieves robust detection in crowded and occluded scenes. Furthermore, we introduce a unified de-duplication framework to eliminate cross-frame redundant counting. To facilitate future research, we also release a new RGB-D Video Object Counting dataset featuring depth information and multiple object categories persequence. Extensive experiments demonstrate that our method achieves a 62.01\% reduction in MAE compared to existing baselines, and also produces consistent improvements in RMSE. We provide the source code at https://github.com/streamer-AP/DG-Net and the dataset at https://huggingface.co/datasets/aerospace123/RGBD-VideoCount.

补充信息

↑