arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于模式对齐的运动学知识图谱:多模态步态分析中的结构化潜在表示学习

Kinematic Knowledge Maps for Pattern Alignment: Structured Latent Representational Learning in Multimodal Gait Analysis

Chen Dong, He Zonglin, Cheung Kenneth M. C

arXiv 2608.20969首次发表:更新:

AI 中文总结

本研究提出ScoliDetect框架,通过运动学知识图谱(KKM)实现多模态步态分析,其介导的融合结合三模态对比预训练提升了脊柱侧凸筛查的泛化性与可解释性,外部ROC-AUC达0.972。

AI 中文摘要

多模态临床人工智能受限于输入对齐不足及缺乏领域特定的可解释表示,尤其在从密集视频流、结构化时间序列和基于模板的运动学文本中学习时。本文提出ScoliDetect,一种基于单目步态视频的青少年特发性脊柱侧凸筛查可解释框架,围绕运动学知识图谱(KKM)及来自序列姿态统计的互补基于模板的运动学文本构建。KKM是固定索引的结构化表示,编码绝对运动、自身骨骼配置及关节-关节信号关联的步态特征,提供锚点参考的多模态融合与因子级解释。我们通过带潜在瓶颈聚合的双向交叉注意力整合视频、KKM及基于模板的运动学文本。在多中心队列(排除后n=1858)中,外部筛查队列的预先指定监督消融实验显示,KKM介导的多模态融合优于单模态模型及后期拼接。在分阶段训练协议下,架构选择后应用三模态对比预训练作为表示初始化,使外部ROC-AUC从0.961提升至0.972。此外,KKM的结构化性质提供固有因子级归因,直接映射至特定运动学阶段及骨骼索引,具备可验证的可解释性。结果表明,将显式结构拓扑嵌入潜在空间可显著提升多模态模式分析系统的泛化性与可解释性。

英文摘要

Multimodal clinical AI is limited by weakly aligned inputs and the absence of domain-specific interpretable representations, particularly when learning from dense video stream, structured time-series, and template-based kinematic text. Here we present ScoliDetect, an explainable framework for adolescent idiopathic scoliosis screening from monocular gait video, built around a kinematic knowledge map (KKM) and complementary template-based kinematic text derived from per-sequence pose statics. KKM is a fixed-index structured representation that encodes gait features across absolute motion, self-skeleton configuration and joint-joint signal correlation, providing anchor-referenced multimodal fusion and factor-level interpretation. We integrate video, KKM, and template-based kinematic text through bidirectional cross-attention with latent-bottleneck aggregation. In a multicenter cohort (n = 1,858 after exclusions), prespecified supervised ablations on an external screening cohort show that KKM-mediated multimodal fusion outperforms unimodal models and late concatenation. Under a staged training protocol, trimodal contrastive pretraining is applied after architecture selection as representation initialization, improving external ROC-AUC from 0.961 to 0.972. Furthermore, the structured nature of the KKM provides inherent, factor-level attributions mapped directly to specific kinematic phases and skeletal indices, offering verifiable interpretability. The results demonstrate that embedding explicit structural topologies into latent spaces significantly enhances both the generalization and explainability of multimodal pattern analysis systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑