arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考单细胞基础模型的类别不平衡问题:跨架构与长尾损失函数的系统性基准测试

Rethinking Class Imbalance for Single-Cell Foundation Models: A Systematic Benchmark Across Architectures and Long-Tail Loss Functions

Zeyu Dong, Jiahui Zhong

arXiv 2609.23325首次发表:更新:

AI 中文总结

本研究系统基准测试了六种长尾损失函数在三种单细胞基础模型上的类别不平衡问题,发现稀有类别失败分为可恢复与不可恢复两种机制,并提供了选择损失函数的实用指南。

AI 中文摘要

单细胞基础模型(scGPT、scBERT、Geneformer)在我们的实验中实现了高达97.5%的细胞类型分类准确率,然而这一总体准确率可能掩盖了对稀有(通常与疾病相关)细胞群体的系统性失败,而长尾损失函数被广泛认为可以解决这一问题。我们提出了一个系统性基准测试,涵盖六种长尾损失函数(交叉熵、加权交叉熵、类别平衡损失、焦点损失、LDAM、对数几率调整softmax),跨越三种架构和三个数据集(多发性硬化症、Zheng68K、人类胰腺),总计162次受控训练运行(3个骨干网络×3个数据集×6种损失函数×3个随机种子)。在普通交叉熵下,总体准确率、宏F1分数和稀有类别召回率之间的差距在所有九种(架构,数据集)设置中保持一致,这一差距由数据集结构而非预训练驱动。稀有类别失败本身分为两种具有不同嵌入几何特征的机制,且在任何损失函数选择之前即可观察到:某些类别可通过合适的损失函数恢复,而其他类别虽然保持线性可分性,但在所有评估的损失函数和架构下被吸收到无关类别的邻域中。在可恢复的类别中,重新加权的有效性由类别的绝对训练集大小预测,而非其在数据集中的占比或数据集的整体不平衡比率。类别平衡损失和LDAM是所有九种设置中最一致的选择,而对数几率调整则以稀有类别的精确率换取召回率,而非同时改善两者。我们的结果既提供了可复用的基准测试,也为将基础模型与不平衡生物数据结合提供了基于机制的实用指南。

英文摘要

Single-cell foundation models (scGPT, scBERT, Geneformer) achieve cell-type classification accuracy up to 97.5% in our experiments, yet this aggregate accuracy can mask systematic failure on rare, often disease-relevant cell populations that long-tail loss functions are widely assumed to address. We present a systematic benchmark of six long-tail loss functions (cross-entropy, weighted CE, class-balanced loss, focal loss, LDAM, logit-adjusted softmax) across three architectures and three datasets (Multiple Sclerosis, Zheng68K, human Pancreas), totaling 162 controlled training runs (3 backbones x 3 datasets x 6 losses x 3 seeds). The gap between overall accuracy, Macro-F1, and rare-class recall under plain cross-entropy is consistent across all nine (architecture, dataset) settings, driven by dataset structure rather than pretraining. Rare-class failure itself splits into two regimes with distinct embedding-geometry signatures, visible before any loss is chosen: some classes are recoverable by the right loss, while others retain linear separability yet are absorbed into unrelated classes' neighborhoods under every evaluated loss and architecture. Among the recoverable classes, the efficacy of reweighting is predicted by a class's absolute training-set size, rather than its share of the dataset or the dataset's overall imbalance ratio. Class-balanced loss and LDAM are the most consistent choices across all nine settings, while logit adjustment trades rare-class precision for recall rather than improving both. Our results give both a reusable benchmark and mechanism-grounded practical guidelines for combining foundation models with imbalanced biological data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑