基于基础模型的联合多动性运动障碍多标签表型分析
Foundation-model-based multi-label phenotyping of combined hyperkinetic movement disorders
- Service of Neurology, Department of Clinical Neurosciences, Lausanne University Hospital (CHUV) and University of Lausanne (UNIL)(洛桑大学医院(CHUV)和洛桑大学)
- Institut du Neurone(神经元研究所)
- Department of Neurosurgery, Military University Hospital of Sfax(苏法军事大学医院神经外科)
- Department of Neurology, Clinique Beau Soleil, Institut Mutualiste Montpelliérain(蒙彼利埃互助研究所博苏莱恩诊所神经内科)
- Movement Disorders Unit, Pediatric Neurology Department, Institut de Recerca, Hospital Sant Joan de Déu(圣胡安·德·迪乌医院儿科神经内科运动障碍单元研究所以)
- European Reference Network for Rare Neurological Diseases (ERN-RND)(欧洲罕见神经系统疾病参考网络(ERN-RND))
- U-703 Centre for Biomedical Research on Rare Diseases (CIBER-ER), Instituto de Salud Carlos III(卡洛斯三世卫生研究所罕见病生物医学研究中心(CIBER-ER))
- Edinburgh Medical School, University of Edinburgh(爱丁堡大学爱丁堡医学院)
- Department of Neurology, CHU Montpellier(蒙彼利埃大学医院神经内科)
- Department of Clinical Neuroscience, Umeå University(于默奥大学临床神经科学系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究将SAM 3与TabICLv2基础模型结合,实现跨年龄和场景的联合多动性运动障碍多标签表型分析,经轻量校准后在不同数据集上获得高稳健性和零假阳性。
AI中文摘要:
运动障碍(MDs)经常共存,然而现象学和严重程度评估显示出显著的评估者间变异性。无标记视频可以提高可重复性,但先前的工作大多局限于单一症状,依赖于标准化采集,并且缺乏跨年龄和跨地点的验证与迁移。我们将两个基础模型组合成一个冻结骨干网络:Segment Anything Model 3(SAM 3)用于密集的逐帧无标记分割,并总结为几何、轮廓和网格运动学信号;以及TabICLv2,一个表格基础模型,用于八种多动性运动障碍现象学的上下文内多标签分类。在21名成人和4名对照的标准化录音上训练后,该模型未经改变地迁移到两个独立数据集:儿科(n=12)和震颤为主的成人(n=20),使用CODY-SAMP量表进行评估;仅患者级别的决策步骤按地点进行了重新校准。在临床医生共识标签下,两个数据集的假阳性均降至零。肌张力障碍完全恢复(儿科7/7;成人留出集15/15),儿童舞蹈病完全恢复(3/3),而震颤在成人中(11/15)通过仅重新校准在富含震颤的队列使其可评估后得以恢复。逐区域效应量分析提供了临床一致、现象学特异性的信号,并确定肌阵挛为主要失败点。与YOLOv8稀疏关键点相比,密集表示在临床医生许可标签下匹配(Jaccard 0.63对0.63),并在临床医生标签共识下显著更稳健(0.93对0.76)。这种冻结的基础模型骨干网络配合轻量级按地点校准,可在不同年龄和从标准化到常规视频中,对共存的运动障碍产生可迁移、可解释、保守的多标签表型分析,在高置信度、临床医生同意的标签上增加了稳健性。在临床使用前需要进行前瞻性多中心验证。
英文摘要:
Movement disorders (MDs) frequently co-occur, yet phenomenological and severity assessment shows substantial inter-rater variability. Markerless video could improve reproducibility, but prior work is largely single-symptom, depends on standardized acquisition, and lacks validation and transfer across ages and sites. We combined two foundation models into one frozen backbone: Segment Anything Model 3 (SAM 3) for dense, per-frame markerless segmentation summarized into geometric, contour and grid kinematic signals, and TabICLv2, a tabular foundation model, for in-context multi-label classification of eight hyperkinetic MD phenomenologies. Trained on standardized recordings of 21 adults and 4 controls, it transferred unchanged to two independent datasets, pediatric (n=12) and tremor-dominant adult (n=20), assessed with the CODY-SAMP scale; only the patient-level decision step was recalibrated per site. Under clinician consensus labels, false positives fell to zero in both datasets. Dystonia recovered perfectly (7/7 pediatric; 15/15 adult held-out), chorea fully in children (3/3), and tremor was recovered in adults (11/15) once a tremor-rich cohort made it evaluable, through recalibration alone. Per-region effect-size analysis gave clinically coherent, phenomenology-specific signals and identified myoclonus as the principal failure. Against YOLOv8 sparse keypoints, the dense representation matched under clinician permissive labels (Jaccard 0.63 vs 0.63) and was markedly more robust under clinician-label consensus (0.93 vs 0.76). This frozen foundation-model backbone with light per-site calibration yields transferable, interpretable, conservative multi-label phenotyping of co-occurring hyperkinetic MDs across ages and from standardized to routine video, adding robustness on high-confidence, clinician-agreed labels. Prospective multi-centre validation is required before clinical use.