arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AuricularWorld:用于CT扫描中耳部结构细粒度分割的分层动作引导世界建模

AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans

Jingwen Yang, Senmao Wang, Luoyao Kang, Runmeng Cui, Keying Zhang, Yunjia Bao, Haifan Gong, Lin Lin, Haiyue Jiang

arXiv 2607.28487首次发表:更新:

发表机构

Plastic Surgery Hospital, Chinese Academy of Medical Sciences & Peking Union Medical College; The Chinese University of Hong Kong; The Chinese University of Hong Kong (Shenzhen)(中国医学科学院北京协和医学院整形外科医院; 香港中文大学; 香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对CT耳部细粒度分割难题,提出分层动作引导的AuricularWorld世界建模框架,通过迭代解剖学推理提升分割精度,使HD95降低超43%,验证了潜在世界模型推理的有效性。

AI 中文摘要

CT图像中耳部结构的细粒度分割颇具挑战性,原因在于耳部在图像中占比小、软骨边界极不规则,且软骨与周围软组织的界面常模糊不清。临床标注可能同时包含含软骨与相邻皮肤的复合结构,以及对应的仅软骨区域,形成嵌套且重叠的标签。我们提出一种基于世界模型的分割框架,该框架支持超越传统前馈预测的迭代解剖学推理。框架构建于编码器-解码器架构之上,在中间潜在空间引入确定性循环状态空间模型,融合多尺度编码器特征与部分解码表示,形成初始化潜在动态的结构观测。推理阶段,模型在无真实标签引导下执行三步潜在滚动,分层解剖学动作更新循环状态并逐步优化潜在表示,生成的潜在轨迹被投影回解码器,与高分辨率特征结合得到最终分割结果。为学习可靠的潜在转换,我们引入平衡分层动作目标,以解决前景稀疏、解剖组缺失及添加与移除操作间的不平衡问题。大量实验表明,所提框架在CT图像中小、不规则、重叠的耳部结构分割中,持续提升分割精度并使HD95降低超过43%,这些结果证明了潜在世界模型推理在挑战性医学图像分割任务中的有效性。

英文摘要

Fine-grained segmentation of auricular structures in CT is challenging because the ear occupies a small image region, cartilage boundaries are highly irregular, and interfaces between cartilage and surrounding soft tissues are often ambiguous. Clinical annotations may also include both composite structures containing cartilage and adjacent skin and their corresponding cartilage-only regions, producing nested and overlapping labels. We propose a world-model-based segmentation framework that enables iterative anatomical reasoning beyond conventional feed-forward prediction. Built on an encoder-decoder architecture, the framework introduces a deterministic recurrent state-space model into the intermediate latent space. Multi-scale encoder features and partially decoded representations are fused to form a structural observation that initializes the latent dynamics. During inference, the model performs a three-step latent rollout without ground-truth guidance. Hierarchical anatomical actions update the recurrent state and progressively refine the latent representation. The resulting latent trajectory is projected back into the decoder and combined with high-resolution features to produce the final segmentation. To learn reliable latent transitions, we introduce a balanced hierarchical action objective that addresses foreground sparsity, missing anatomical groups, and imbalance between add and remove operations. Extensive experiments show that the proposed framework consistently improves segmentation accuracy and reduces HD95 by more than 43% for small, irregular, and overlapping auricular structures in CT. These results demonstrate the effectiveness of latent world-model reasoning for challenging medical image segmentation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑