发表机构
CSIRO(澳大利亚联邦科学与工业研究组织)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
InfoTaxa通过信息校准分析,诊断出细粒度无标签聚类在物种层面的瓶颈既源于聚类方法也源于图像表示,并利用DNA作为审计信号量化了信息差距。
AI 中文摘要
对冻结的预训练视觉嵌入进行无标签聚类,为生物多样性监测提供了一条可扩展的途径,但仅基于图像的细粒度分类学表现出一种一致的由粗到细的失败模式:聚类能够恢复大致的分类学结构,但在物种层面趋于平稳。我们通过信息校准的聚类分析,在BIOSCAN-5M上研究了这一行为。使用UMAP和HDBSCAN的BioCLIP~2特征在科级别达到$0.79$的AMI,在属级别达到$0.67$,显著优于先前的图像基线,并与基于oracle-$K$、基于图以及学习型聚类头在相同冻结特征上的表现相当。为了诊断剩余的瓶颈是方法限制还是信息限制,我们引入了InfoTaxa,它将聚类效率——即无监督划分恢复的探针估计图像信息比例——与配对的DNA作为仅用于审计的信号(而非推理输入)相结合。密度管线在目和科级别分别恢复了约$0.90$和$0.81$的图像可用信息。留出的晚期融合探针显示,将DNA添加到图像嵌入中可将物种级别的预测误差减少约两个比特。鲁棒性分析涵盖了多种图像编码器、已描述物种和稀有类子集、探针诊断以及留出物种的粗粒度泛化和同物种检索。因此,在测试的设置中,物种级别的无标签聚类既受聚类限制又受表示限制:改进聚类可能恢复额外的图像暴露结构,但无法单独弥合DNA审计的信息差距。
英文摘要
Label-free clustering of frozen pretrained visual embeddings offers a scalable route to biodiversity monitoring, but image-only fine-grained taxonomy exhibits a consistent coarse-to-fine failure mode: clusters recover broad taxonomic structure yet plateau at species level. We study this behaviour on BIOSCAN-5M through an information-calibrated clustering analysis. BioCLIP~2 features with UMAP and HDBSCAN reach $0.79$ AMI at family and $0.67$ at genus, substantially improving over the prior image baseline and remaining competitive with oracle-$K$, graph-based, and learned clustering heads on the same frozen features. To diagnose whether the remaining plateau is method-limited or information-limited, we introduce InfoTaxa, which combines clustering efficiency---the fraction of probe-estimated image information recovered by an unsupervised partition---with paired DNA as an audit signal only, not an inference input. The density pipeline recovers approximately $0.90$ and $0.81$ of the image-available information at order and family, respectively. Held-out late-fusion probes show that adding DNA to the image embedding reduces species-level prediction error by approximately two bits. Robustness analyses cover multiple image encoders, described-species and rare-class subsets, probe diagnostics, and held-out-species coarse-rank generalisation and same-species retrieval. Thus, in the tested setting, species-level label-free clustering is both clustering-limited and representation-limited: improved clustering may recover additional image-exposed structure, but cannot close the DNA-audited information gap alone.