arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SlideMix:通过多模态混洗增强全切片图像分析

SlideMix: Enhancing Whole Slide Image Analysis via Multimodal Shuffling

Chad Wong, Sicheng Chen, Tianyi Zhang, Enhui Chai, Yueming Jin, Zeyu Liu, Fei Xia

arXiv 2609.00396首次发表:更新:

发表机构

University of California, Irvine; National University of Singapore; PuzzleLogic Pte Ltd(加利福尼亚大学欧文分校; 新加坡国立大学; PuzzleLogic私人有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SlideMix是适用于多实例学习全切片图像分析的模型无关多模态增强框架,通过检索增强VLM区域选择、原位图块混洗等提升诊断准确率与泛化能力,在11个WSI数据集上表现优于基线。

AI 中文摘要

组织病理学全切片图像(WSI)是癌症诊断的核心,但它们的千兆像素级规模、组织异质性、弱切片级监督、稀疏诊断区域及多尺度证据,使得稳健的自动化分析颇具挑战。多实例学习(MIL)被广泛用于将图块级特征聚合为切片级预测,然而现有的增强策略常常扰动组织区域,却未保留诊断相关性、切片上下文或跨尺度结构。我们提出SlideMix,一种适用于基于MIL的WSI分析的模型无关多模态增强框架。SlideMix采用基于检索增强视觉语言模型(VLM)的视觉语言自适应区域选择器,识别诊断相关区域并减少弱标签噪声;随后在有意义的组织区域内执行原位图块混洗,以混合特征嵌入同时保留切片级上下文;基于VLM的软标签模块对混合样本进行监督,同时多因素、损失驱动的在线课程学习反馈方案自适应控制混洗粒度、特征相似度及混洗比例,以促进跨尺度表示学习。在包含20523张切片、8项诊断任务及10种WSI骨干的11个WSI数据集上,SlideMix在多数场景下提升了准确率与泛化能力,且与已有的增强基线相比表现更优,为构建更稳健、可扩展的数字病理学模型提供了一种简单的即插即用方法。源代码:this https URL

英文摘要

Histopathological whole slide images (WSIs) are central to cancer diagnosis, but their gigapixel scale, tissue heterogeneity, weak slide-level supervision, sparse diagnostic regions, and multi-scale evidence make robust automated analysis challenging. Multiple instance learning (MIL) is widely used to aggregate tile-level features into slide-level predictions, yet existing augmentation strategies often perturb tissue regions without preserving diagnostic relevance, slide context, or cross-scale structure. We propose SlideMix, a model-agnostic multimodal augmentation framework for MIL-based WSI analysis. SlideMix uses a retrieval-augmented vision-language model (VLM)-based Visual-Language Adaptive Region selector to identify diagnostically relevant regions and reduce weak-label noise. It then performs In-place Tile Shuffling within meaningful tissue regions to mix feature embeddings while preserving slide-level context. A VLM-based soft-labeling module supervises mixed samples, while a multi-factor, loss-driven online Curriculum-Learning Feedback scheme adaptively controls shuffle granularity, feature similarity, and shuffle ratio to promote cross-scale representation learning. Across 11 WSI datasets comprising 20,523 slides, 8 diagnostic tasks, and 10 WSI backbones, SlideMix improves accuracy and generalization in most settings and compares favorably with established augmentation baselines, providing a simple plug-and-play approach for more robust and scalable digital pathology models. Source code: https://github.com/Xia-Research-Lab/SlideMix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑