AI 中文总结
针对动物姿态跟踪的跨物种泛化与精度问题,提出有监督、无监督两种基于视觉基础模型的方法,在APTv2和TigDog数据集上实现了精度与泛化性的平衡,为动物保护应用提供实用方案。
AI 中文摘要
动物姿态估计与跟踪对野生动物监测和保护研究至关重要,由于专家标注时间有限,自动化方法必不可少。人类姿态估计与跟踪因大量标注数据集已取得快速进展,但动物姿态因物种间形态和行为差异大、标注数据有限而颇具挑战性。现有方法要么从标注数据集(如APTv2)优化通用关键点定位,泛化性差;要么用视觉跟踪跟踪自定义关键点,性能受损。本文证明,基于大数据训练的视觉基础模型可在有限标注数据下有效用于动物姿态跟踪。我们提出两种模型,一种无监督,一种有监督,用于跟踪视频中用户选定的关键点。有监督方法采用关键点提示编码器,将参考帧的结构先验显式注入特征匹配,实现更优跟踪精度。无监督路径则利用多样化基础模型特征进行无训练对应匹配,提供强跨物种鲁棒性。在具有挑战性的动物视频基准数据集APTv2和TigDog上的广泛评估表明,我们的框架在精度与泛化性间保持有效平衡,表现出色,为实际动物行为分析和保护应用提供了实用解决方案。
英文摘要
Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated approaches are imperative. While human pose estimation and tracking has seen rapid progress thanks to large annotated datasets, animal pose remain challenging, due to large morphological and behavioural differences between species and limited annotated data. Existing approaches either optimise generic keypoint localisation from annotated datasets (such as APTv2) with poor generalisation, or track custom keypoints using visual tracking, at the cost of performance. In this paper, we demonstrate that vision foundation models trained on large datasets can be used effectively to track animal pose with limited labelled data. We propose two models, one unsupervised and the other supervised, to track user-selected keypoints in videos. The supervised approach delivers superior tracking accuracy by employing a keypoint prompt encoder to explicitly inject structural priors from a reference frame into feature matching. In parallel, the unsupervised route provides strong cross-species robustness by leveraging diverse foundation-model features for training-free correspondence matching. Extensive evaluation on challenging animal video benchmarks APTv2 and TigDog demonstrates that our framework achieves strong performance while maintaining an effective balance between accuracy and generalisation, offering a practical solution for real-world animal behaviour analysis and conservation applications.
CommentsAccepted for presentation at the ECCV 2026 Workshop on CV4Ecology