arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VastMAT:大规模多类别多动物跟踪基准

VastMAT: A Large-Scale Multi-Category Benchmark for Multi-Animal Tracking

Zhizhen Li, Zan Wang, Huidong Peng, Bohan Tan, Shimin Shan, Yu Liu, Liang Peng

arXiv 2609.34390首次发表:更新:

发表机构

Dalian University of Technology; Wuhan University; University of North Texas; The Hong Kong University of Science and Technology(大连理工大学; 武汉大学; 北德克萨斯大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VastMAT是一个大规模多类别多动物跟踪基准,包含2947个视频、337个类别和366万多个边界框,并提出了CDA模块,在不训练的情况下提升跟踪性能,解决低重叠关联挑战。

AI 中文摘要

多动物跟踪(MAT)支持对动物运动、行为和群体互动的研究。然而,通用的多目标跟踪(MOT)基准主要关注行人和车辆,而专门的MAT基准在同时支持广泛的动物覆盖、大规模视频数据和视频内多实例关联方面仍然有限。为弥补这一空白,我们提出了VastMAT,它具有四个关键特征:(1)大规模。它包含2,947个视频,共1,002,562个标注帧,总计27.85小时。(2)广泛的类别覆盖。这些视频涵盖337个动物类别,具有多样的形态和运动模式。(3)广泛的实例标注。它提供了3,663,248个边界框和22,883条身份轨迹——据我们所知,这两项数据在专门的MAT基准中都是最多的。(4)高质量标注。为确保可靠性,标注经过迭代的专家审查和修正,并通过独立的重新标注审计来评估质量。为了系统评估跟踪性能和跨类别泛化能力,我们建立了已见类别和类别不相交的未见类别协议,并在两种协议下评估了八种代表性的MOT方法。在这些协议下,最高的基线HOTA分数分别为66.37%和52.90%,凸显了跟踪未见动物的挑战。为了解决我们分析中揭示的低重叠关联挑战,我们提出了中心距离增强关联(CDA),这是一种轻量级模块,自适应地将IoU与由框自身尺度归一化的中心相似度相结合。无需额外训练,CDA在两种协议下分别将TrackTrack的HOTA提高了1.58和1.31个百分点。为促进进一步的MAT研究,我们将公开发布我们的基准和代码。

英文摘要

Multi-animal tracking (MAT) supports the study of animal movement, behavior, and group interactions. However, general multi-object tracking (MOT) benchmarks primarily focus on pedestrians and vehicles, whereas dedicated MAT benchmarks remain limited in jointly supporting broad animal coverage, large-scale video data, and extensive within-video multi-instance association. To address this gap, we introduce VastMAT, which has four key characteristics: (1) Large scale. It comprises 2,947 videos with 1,002,562 annotated frames, totaling 27.85 hours. (2) Broad category coverage. These videos cover 337 animal categories with diverse morphologies and motion patterns. (3) Extensive instance annotations. It provides 3,663,248 bounding boxes and 22,883 identity trajectories---to our knowledge, the largest numbers of both among dedicated MAT benchmarks. (4) High-quality annotations. To ensure reliability, annotations undergo iterative expert review and correction, and quality is assessed through an independent reannotation audit. To systematically assess tracking performance and cross-category generalization, we establish Seen-category and category-disjoint Unseen-category protocols, and evaluate eight representative MOT methods under both protocols. Under these protocols, the highest baseline HOTA scores are 66.37\% and 52.90\%, respectively, highlighting the challenge of tracking unseen animals. To address the low-overlap association challenge revealed by our analysis, we propose Center-Distance-Augmented Association (CDA), a lightweight module that adaptively combines IoU with center similarity normalized by the boxes' own scales. Without additional training, CDA improves TrackTrack's HOTA by 1.58 and 1.31 percentage points under the two protocols, respectively. To facilitate further MAT research, we will publicly release our benchmark and code.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑