发表机构
Virginia Tech; Accenture(弗吉尼亚理工大学; 埃森哲)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究审计15种冻结编码器在采集分布偏移下的鲁棒性,发现跨域性能骤降,提出CBR方法改善部分场景,强调血液学FM基准需同时审计准确率、校准等多维度指标。
AI 中文摘要
冻结的血液学基础模型(FM)嵌入在域内白细胞(WBC)准确率接近饱和,但临床部署要求其在不同扫描仪、站点、染色和制备流程下具备可靠性。我们在四个公开单细胞采集域中,针对15种冻结编码器(血液学、病理学及通用视觉),从准确率鲁棒性和校准两个维度进行审计。域内线性探测宏F1值处于饱和状态(0.98-0.997),但跨数据集宏F1值下降34%-72%且排名发生变化:在基准共享的224像素输入下,域内表现最佳的DinoBloom-L在偏移最大的目标数据集MLL23上,排名从第1降至第15中的第10位,落后于RedDino及若干通用和病理学编码器。排名转移依赖探测方式:平均而言,1-NN检索比源域适配的线性头更稳定(中位数ρ为0.65 vs 0.45),但两种探测方式均无法通用预测目标域鲁棒性。校准也出现崩溃:源域训练的探测在域内几乎校准(预期校准误差ECE为0.004),但在跨域时置信错误(ECE为0.35),且源域适配的温度缩放迁移效果差。我们进一步审计预训练暴露情况,发现MLL23是DinoBloom的内部队列;由于DinoBloom唯一的保留数据集也是我们的源域,该基准无法将暴露与扫描仪相关偏移隔离开。无标签自适应和基于边际熵的模型选择在平衡评估下表现安全,但在现实的WBC类别先验偏移下失效。类别平衡重标准化(CBR)是一种无训练的伪标签平衡特征归一化方法,可改善所有评估目标先验场景的均值并部分改善校准,但仍存在编码器级例外和残留校准误差。因此,血液学FM基准必须同时审计准确率、校准、暴露和类别先验鲁棒性。
英文摘要
Frozen hematology foundation-model (FM) embeddings reach near-saturated in-domain white-blood-cell (WBC) accuracy, but clinical deployment demands reliability across scanners, sites, stains and preparation pipelines. We audit 15 frozen encoders (hematology, pathology, and general vision) across four public single-cell acquisition domains along two axes: accuracy robustness and calibration. In-domain linear-probe macro-F1 is saturated (0.98-0.997), yet cross-dataset macro-F1 drops 34-72% and rankings re-order: DinoBloom-L, the in-domain best, falls to 10th of 15 on the most-shifted target (MLL23) at the benchmark's shared 224-px input, behind RedDino and several general and pathology encoders. Rank transfer is probe-dependent: 1-NN retrieval is more stable on average than a source-fitted linear head (median $ρ$ 0.65 vs 0.45), but neither probe universally predicts target robustness. Calibration also collapses: source-trained probes are nearly calibrated in-domain (expected calibration error, ECE, 0.004) but confidently wrong off-domain (ECE 0.35), and source-fitted temperature scaling transfers poorly. We further audit pretraining exposure and identify MLL23 as DinoBloom's internal cohort; because DinoBloom's only held-out dataset is also our source domain, this benchmark cannot isolate exposure from scanner-associated shift. Label-free adaptation and marginal-entropy-based model selection appear safe under balanced evaluation but fail under realistic WBC class-prior shift. Class-Balanced Re-standardization (CBR), a training-free pseudo-label-balanced feature normalization, improves all evaluated target-prior scenario means and partially improves calibration, although encoder-level exceptions and residual miscalibration remain. Hematology FM benchmarks must therefore jointly audit accuracy, calibration, exposure, and class-prior robustness.
CommentsAccepted at the HemaRAI 2026 workshop (MICCAI 2026 satellite event), oral presentation; to appear in MICCAI 2026 Satellite Events, LNCS, Springer. 25 pages (10 main incl. references + 15 supplementary), 4 figures. Project page: https://jaishrm07.github.io/hematology-fm-robustness/