动物再识别学到了什么?线性生物学概念及其在视觉表示中的起源
What Does Animal Re-Identification Learn? Linear Biological Concepts and Their Origins in Visual Representations
- Hasso-Plattner Institute(哈索·普拉特纳研究院)
- University of Potsdam(波茨坦大学)
- Fraunhofer Heinrich Hertz Institute(弗劳恩霍夫海因里希·赫兹研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过可解释性方法揭示,基于ViT的动物再识别模型在无显式监督下,其表示中涌现出可泛化的性别和年龄线性概念,且训练仅重定位而非创造这些概念,为野生动物监测的可审计性提供了新视角。
AI中文摘要:
保护工作日益依赖相机陷阱,这些设备收集的野生动物图像远超专家人工分析的容量,使得动物再识别(Re-ID)对于监测个体和种群至关重要。然而,对于基于ViT的再识别模型,理解哪些线索驱动模型决策颇具挑战性,因为其度量学习目标并未为生物学概念提供显式监督。我们探究这类模型是否仍能沿生物学意义的轴组织其表示。使用针对西部低地大猩猩再识别、以三元组边际损失微调的DINOv3骨干网络,我们发现性别和年龄作为线性方向涌现,并能泛化到未见的个体,AUROC最高达0.91,且可从每个个体单张图像中恢复。激活引导进一步表明,性别方向被模型因果性地使用,可将显著比例的预测翻转为相反性别。对比现成与微调骨干网络显示,再识别训练并未创造这些概念,而是将其在网络中重新定位。最后,数据归因揭示,我们发现的表示反映了一个分级的生物学轴,在种群中被冗余编码,并受视觉上模糊的个体塑造。综合来看,这些发现展示了可解释性如何揭示再识别表示的生物学结构和失败模式,为野生动物监测的可审计计算机视觉迈出一步。
英文摘要:
Conservation increasingly relies on camera traps that collect more wildlife imagery than experts can manually analyze, making animal re-identification (Re-ID) essential for monitoring individuals and populations. Yet understanding which cues drive model decisions is challenging for ViT-based Re-ID models, whose metric-learning objectives provide no explicit supervision for biological concepts. We ask whether such models nonetheless organize their representations along biologically meaningful axes. Using a DINOv3 backbone fine-tuned for Western lowland gorilla Re-ID with triplet-margin loss, we find that sex and age emerge as linear directions that generalize to held-out individuals, reaching up to 0.91 AUROC and being recoverable from a single image per individual. Activation steering further shows that the sex direction is causally used by the model, flipping a significant fraction of predictions to the opposite sex. Comparing off-the-shelf and fine-tuned backbones shows that Re-ID training does not create these concepts, but relocates them across the network. Finally, data attribution reveals that the representation we find reflects a graded biological axis, is redundantly encoded across the population and shaped by visually ambiguous individuals. Together, these findings show how interpretability can uncover both the biological structure and failure modes of Re-ID representations, providing a step toward auditable computer vision for wildlife monitoring.