arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结构感知的吞咽荧光透视检查关键点定位

Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study

Kai Zhou, Chuanshen Chen, Runhao Zeng, Meng Dai, Yifan Yang, Jinwu Hu, Daiyuan Li, Mingkui Tan, Fei Liu

arXiv 2610.07726首次发表:更新:

发表机构

South China University of Technology; Shenzhen MSU-BIT University; The Third Affiliated Hospital of Sun Yat-sen University; Electric Power Research Institute, China South Grid(华南理工大学; 深圳北理莫斯科大学; 中山大学附属第三医院; 中国南方电网电力科学研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VFSS关键点定位中空间偏差和数据效率问题,提出VFSSKep数据集和S$^3$KL框架,通过结构感知学习与一致性策略,在半监督下超越全监督性能。

AI 中文摘要

吞咽荧光透视检查(VFSS)是诊断吞咽障碍的金标准之一,提供吞咽过程的动态X射线成像。VFSS中的自动化运动学分析从根本上依赖于精确的解剖关键点定位。然而,现有研究集中于有限的关键点(如颈椎或舌骨),忽略了软腭等关键区域,且仅标注活跃吞咽片段,忽略了大量非吞咽数据,导致数据效率低下。此外,通过标准半监督学习利用这些未标注数据是次优的,因为通用方法容易产生空间偏差。在布局固定的医学X射线中,模型倾向于记忆绝对坐标而非理解解剖结构。为应对这些挑战,我们引入了VFSSKep,一个新数据集,将标注扩展到软腭并纳入大规模未标注数据。我们进一步提出了S$^3$KL,一种结构感知的半监督关键点定位框架,旨在克服空间偏差。它集成了结构感知学习策略以提取高分辨率结构线索用于结构感知表示学习,以及带有块打乱的结构表示一致性学习策略以强制不变的结构识别。实验表明,我们的方法达到了最先进的半监督性能,即使仅使用未标注数据和25%的标注数据,也超过了使用100%标注数据的全监督学习。代码和数据将在以下网址公开:this https URL。

英文摘要

Videofluoroscopic Swallowing Study (VFSS) is one of the gold standard for diagnosing swallowing disorders, providing dynamic X-ray imaging of the swallowing process. Automated kinematic analysis in VFSS relies fundamentally on precise anatomical keypoint localization. However, existing studies focus on limited keypoints (e.g., cervical vertebrae or the hyoid) and overlook critical regions such as the soft palate, while annotating only active swallowing segments and ignoring abundant non-swallowing data, resulting in poor data efficiency. Moreover, leveraging this unlabeled data via standard semi-supervised learning is suboptimal, as generic methods are prone to spatial bias. In medical X-rays with fixed layouts, models tend to memorize absolute coordinates rather than understanding anatomical structures. To tackle these challenges, we introduce VFSSKep, a novel dataset that extends annotations to the soft palate and incorporates large-scale unlabeled data. We further propose S$^3$KL, a Structure-aware Semi-Supervised Keypoint Localization framework designed to overcome spatial bias. It integrates a Structure-Aware Learning strategy to extract high-resolution structural cues for structure-aware representation learning, and a Structural Representation Consistency Learning strategy with block shuffling to enforce invariant structural recognition. Experiments show our method achieves state-of-the-art semi-supervised performance, even with unlabeled and 25% labeled data surpassing fully supervised learning with 100% labeled data. Code and data will be made publicly available at: https://github.com/kaai520/S3KL.

CommentsAccepted by ICME 2026 Oral

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑