arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

冻结步态在空间遮挡下的预测:一种IMU监督的跨模态蒸馏方法

Freezing of Gait Prediction Under Spatial Occlusion: An IMU-Supervised Cross-Modal Distillation Approach

Chandan Biswas, Aryan Singh, Anabik Pal

arXiv 2609.09826首次发表:更新:

发表机构

NeuroAI Fusion Labs; Indian Institute of Science Education and Research Berhampur(NeuroAI 融合实验室; 印度科学教育与研究学院贝汉普尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视频FOG检测受下肢自遮挡影响的问题,提出跨模态子空间蒸馏框架,利用IMU监督视觉模型,双流融合像素与骨骼特征,在零穿戴下实现高精度预测。

AI 中文摘要

帕金森病是一种进行性神经退行性疾病,其特征是运动控制逐渐恶化。自动冻结步态(FOG)检测支持对步态相关的运动障碍进行客观评估。目前有两种常见方法用于FOG预测:(i)分析患者运动的视频记录,以及(ii)分析附着在患者下肢的惯性测量单元(IMU)可穿戴传感器收集的数据。基于视频的方法在连续原地转身任务中可能遭受检测错误,因为下肢会发生大量的几何自遮挡,从而降低姿态估计的准确性。基于IMU的方法通常受视觉遮挡影响较小;然而,它们难以在临床或实验室环境之外部署,因为传感器必须牢固地附着并在整个评估过程中保持原位。受此启发,我们提出了一种跨模态子空间蒸馏框架,通过结合IMU的准确性与基于视频的实用性来减轻单模态FOG检测的局限性。我们从预训练的运动学“神谕”中提取不变潜在拓扑,以在训练期间结构性地监督非编码视觉架构。为了解决严重空间遮挡时期的问题,双流视觉模型概率性地融合骨骼图节点和连续空间像素,随着关节跟踪置信度下降,动态地将依赖转移到不间断的像素边界。针对帕金森患者执行连续360°转弯的公共多模态序列数据集进行评估,实证结果表明,应用感觉边界拓扑严格减轻了跟踪评估熵。我们的约束优化证实,在零可穿戴推理环境中可以实现高精度的FOG预测界限。

英文摘要

Parkinson's disease is a progressive neurodegenerative disorder characterised by gradual deterioration of movement control. Automated freezing-of-gait (FOG) detection supports the objective assessment of gait-related motor impairment. Two common approaches are used for FOG prediction: (i) analysing video recordings of the patient's movements and (ii) analysing data collected using inertial measurement unit (IMU) wearable sensors attached to the patient's lower limbs. Video-based approaches may suffer detection errors during continuous turning-in-place tasks because the lower limbs undergo substantial geometric self-occlusion, degrading pose-estimation accuracy. IMU-based approaches are generally less affected by visual occlusion; however, they are difficult to deploy outside clinical or laboratory settings, as the sensors must be attached securely and remain in place throughout the assessment. Motivated by this, we propose a cross-modal subspace distillation framework to mitigate the limitations of unimodal FOG detection by combining IMU accuracy with video-based practicality. We extract invariant latent topologies from a pre-trained kinematic oracle to structurally supervise a non-encoded visual architecture during training. To resolve periods of severe spatial occlusion, a dual-stream visual model probabilistically fuses skeletal graph nodes and continuous spatial pixels, dynamically shifting reliance to uninterrupted pixel boundaries as joint tracking confidence drops. Evaluated against a public, multi-modal sequence dataset of Parkinson's individuals executing continuous $360^\circ$ turns, empirical results demonstrate that applying sensory boundary topologies strictly mitigates tracking evaluation entropy. Our constrained optimisation confirms that highly precise FOG prediction bounds can be achieved over zero-wearable inference environments.

Comments9 pages , 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑