arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向可靠的基于AI的组织学染色:对非配对生成模型的缩放与不确定性的系统研究

Towards Reliable AI-Based Histological Staining: A Systematic Study of Scaling and Uncertainty in Unpaired Generative Models

Qasim Siddiqui, Adrian Friebel, Maiju Myllys, Zaynab Hobloss, Daniela Gonzalez, Ahmed Ghallab, Stefan Hoehme

arXiv 2608.24626首次发表:更新:

AI 中文总结

本研究针对首个公开的配对H&E-SR小鼠肝脏数据集,系统基准测试6种无监督图像到图像架构,揭示需联合多维度指标评估虚拟染色,且公开相关资源。

AI 中文摘要

肝纤维化是慢性肝病长期预后的主要预测指标,通过胶原蛋白含量的组织学评估进行分期。天狼星红(SR)是标准定量读数(胶原蛋白面积比例,CPA),但并非每个临床中心都能获取,且会消耗组织、时间和试剂成本,超出常规苏木精-伊红(H&E)染色的范畴。基于AI的虚拟染色可直接从H&E生成SR,但缺乏对无监督模型的系统基准测试,且其预测不确定性未被量化,尽管视觉上看似合理的输出可能无法忠实地再现潜在的组织结构。因此,我们在新发布的配对H&E-SR小鼠肝脏数据集(该任务的首个公开资源)上,针对54种缩放配置对6种无监督图像到图像架构(基于GAN和基于扩散的)进行基准测试。每种配置在感知、分布和任务特定维度以及盲法专家读者研究中进行联合评估;随后将每个类别中表现最佳的模型重新训练为深度集成模型,这是首次对无监督染色间架构的认知不确定性进行系统比较。在不同类别中,感知质量、任务特定误差和集成一致性在很大程度上是模型适用性的独立维度:基于GAN的方法在感知指标上紧密聚类,但在任务误差和集成一致性上差异显著,而基于扩散的方法(CycleDiffusion)在所有三个维度上均存在质的差异。没有单一指标能捕捉这些差异,因此可靠的虚拟染色需要同时报告和选择这三个维度。数据集、分块流程、模型和评估代码均已公开。

英文摘要

Liver fibrosis, the principal predictor of long-term outcome in chronic liver disease, is staged from histological estimates of collagen content. Sirius Red (SR) provides the standard quantitative readout (collagen proportionate area, CPA) but is not acquired at every clinical centre and consumes tissue, time, and reagent cost beyond the routine Hematoxylin and eosin (H&E) stain. AI-based virtual staining can generate SR directly from H&E, yet systematic benchmarks of unsupervised models are scarce and their predictive uncertainty has not been quantified, even though visually plausible outputs may not faithfully reproduce the underlying tissue structure. We therefore benchmark six unsupervised image-to-image architectures (GAN-based and diffusion-based) across 54 scaling configurations on a newly released paired H&E to SR mouse liver dataset, the first open resource for this translation task. Each configuration is evaluated jointly on perceptual, distributional, and task-specific axes plus a blinded expert reader study; the best per family is then retrained as a deep ensemble, the first systematic comparison of epistemic uncertainty across unsupervised stain-to-stain architectures. Across families, perceptual quality, task-specific error, and ensemble agreement measure largely independent axes of model fitness: GAN-based methods cluster tightly on perceptual metrics yet differ substantially on task error and ensemble agreement, while the diffusion-based method (CycleDiffusion) is qualitatively different on all three. No single metric captures these differences, so reliable virtual staining requires reporting and selecting on all three jointly. The dataset, tiling pipeline, models, and evaluation code are released publicly.

Comments37th British Machine Vision Conference, November 2026, Lancaster, UK

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑