arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从空间语义到时间上下文:利用注视轨迹实现弱监督医学图像分割

From Spatial Semantics to Temporal Context: Leveraging Gaze Trajectory for Weakly Supervised Medical Image Segmentation

Shaoxuan Wu, Xiao Zhang, Xiaodi Zhao, Yunzhi Tian, Yilin Tang, Jun Feng

arXiv 2607.26542首次发表:更新:

AI 中文总结

该研究针对医学图像分割依赖像素级标注的问题,提出TrailNet模型,结合注视点与轨迹建模时空上下文,引入循环蒸馏实现无注视推理,在公开数据集上取得优于现有方法的Dice分数。

AI 中文摘要

医学图像分割高度依赖耗时费力的像素级标注,眼动追踪提供了可自然融入临床流程的高性价比解决方案。眼动仪记录的注视数据通过注视点传递临床医生关注的空间区域,通过轨迹传递医生逐步视觉感知的时间上下文。然而,有效建模时间轨迹仍具挑战性,探索性注视导致的眼动噪声极大限制了分割性能。为克服这些局限,我们提出轨迹引导的不确定性感知网络(TrailNet),该网络通过联合利用注视点与轨迹,将注视监督医学图像分割从空间语义建模扩展至时间上下文。具体而言,所提出的轨迹引导时空编码器对时间上下文进行建模,并与图像空间语义建立互补交互以强化目标感知;多尺度不确定性解码器利用类别互斥约束生成确定性预测,缓解噪声引发的监督不确定性。为实现无注视推理,我们进一步引入循环蒸馏策略,通过师生网络传递特征级知识。在两个公开数据集上的实验结果表明,TrailNet优于现有最优方法,分别取得81.25%和81.85%的Dice分数。

英文摘要

Medical image segmentation heavily depends on labor-intensive and time-consuming pixel-level annotations. Eye tracking offers a cost-effective solution that can be naturally integrated into clinical workflows. Recorded by eye trackers, gaze conveys the spatial regions of clinicians' attention through fixations and the temporal context of clinicians' progressive visual perception from trajectories. Nevertheless, effective modeling of temporal trajectories remains challenging, and noise in gaze caused by exploratory fixations greatly limits segmentation performance. To overcome these limitations, we propose the Trajectory-guided Uncertainty-aware Network (TrailNet), which exploits gaze-supervised medical image segmentation from spatial semantics modeling to temporal context by jointly leveraging fixations and trajectories. Specifically, the proposed trajectory-guided spatio-temporal encoder models temporal context and establishes complementary interactions with image spatial semantics to strengthen target perception. Furthermore, the multi-scale uncertainty decoder leverages category mutual-exclusivity constraints to produce deterministic predictions and mitigate supervision uncertainty induced by noise. To enable gaze-free inference, we further introduce a cycle distillation strategy that transfers feature-level knowledge via teacher-student networks. Experimental results on two public datasets demonstrate that TrailNet outperforms state-of-the-art methods, achieving Dice scores of 81.25% and 81.85%, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑