基于深度学习的喉高速视频内镜图像中喉部结构分割
Laryngeal Structure Segmentation in High-Speed Videoendoscopy Using Deep Learning
- Michigan State University(密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究训练U-Net模型分割喉高速视频内镜图像中的喉部结构,应用于持续元音和连贯语音数据,总体准确率超95%,为自动化喉部图像分析及异常检测提供可靠工具。
AI中文摘要:
喉高速视频内镜(HSV)为观察不同喉部结构的运动以及在不同发声条件下声带的振动行为提供了一种有效手段。喉部组织的分割能够分析不同组织结构及其动力学,有助于表征喉部肌肉在发声过程中的参与情况。鉴于HSV帧数量庞大,自动化该任务势在必行。尽管以往研究已采用基于深度学习的方法分割喉部结构,但这些方法尚未应用于连贯语音期间的HSV数据,而此类数据因组织过度运动及光纤图像采集相关的图像质量限制而面临重大挑战。将深度学习应用于连贯语音数据对于捕捉非平稳喉部行为及识别与嗓音障碍相关的异常模式至关重要。本研究旨在通过训练U-Net模型来检测杓会厌襞和杓状软骨、声带、会厌及声门区,以弥补上述空白,所用HSV数据来自正常嗓音和障碍嗓音的持续元音发声及连贯语音。研究采用了包括噪声去除和直方图均衡化在内的图像预处理技术,以提高训练HSV图像质量并增强网络性能。最后,为评估网络的准确性和可靠性,在定性目视检查测试图像的同时,使用了定量性能指标。所开发网络表现出高性能,总体准确率超过95%,确立了其作为自动化喉部图像分析、喉部动力学定量表征以及未来临床环境中异常喉部行为检测的可靠工具的潜力。
英文摘要:
Laryngeal high-speed videoendoscopy (HSV) offers an effective means of observing the motion of different laryngeal structures along with vibratory behaviors of the vocal folds under various voicing conditions. Segmentation of laryngeal tissues enables analysis of different tissue structures and their dynamics, helping characterize the involvement of laryngeal muscles in voice production. Given the large number of HSV frames, automating this task is imperative. While deep learning-based methods have been implemented in previous studies to segment laryngeal structures, they have not been applied to HSV data during connected speech, which poses significant challenges due to excessive tissue movements and image quality limitations associated with fiberoptic image acquisition. The application of deep learning to connected speech data is critical for capturing nonstationary laryngeal behaviors and identifying anomalous patterns associated with voice disorders. The present study aims to address these gaps by training U-Net models to detect the aryepiglottic folds and arytenoid cartilages, vocal folds, epiglottis, and glottal area, using HSV data from both sustained vowel phonation and connected speech obtained from normophonic and disordered voices. Image pre-processing techniques, including noise removal and histogram equalization, were applied to improve the quality of the training HSV images and enhance network performance. Finally, to evaluate the accuracy and reliability of the networks, quantitative performance metrics were used alongside qualitative visual inspection of the test images. The high performance of the developed networks, with overall accuracies exceeding 95%, establishes their potential as reliable tools for automated laryngeal image analysis, quantitative characterization of laryngeal dynamics, and future detection of anomalous laryngeal behaviors in clinical settings.