发表机构
UNIST(蔚山国立科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出语言对齐的运动表示方法,结合 Bi-GRU、文本对齐训练、领域自适应与参数合并,实现领域泛化的 UPDRS-步态严重程度评估,在 MoCha 挑战赛中以 637K 参数取得宏 F1 0.57,排名第三。
AI 中文摘要
在本工作中,我们引入了语言对齐的运动表示,用于领域泛化的 UPDRS-步态严重程度评估,旨在学习语义结构化的运动特征,使其能够泛化到异质的临床领域。我们首先使用 Bi-GRU 骨干网络学习运动表示,该网络捕捉 SMPL 序列的时间动态。在模型训练之前,使用 Qwen2.5-7B-Instruct 离线生成运动描述。随后,骨干网络通过分类目标和文本对齐目标进行训练,以学习具有判别性和语义结构化的运动特征,同时考虑训练数据中存在的类别不平衡问题。接下来,我们将学习到的骨干网络独立地适应每个源领域,使模型能够捕捉特定领域的运动特征。然后,在参数层面合并各个源领域的模型,将跨源领域的互补知识整合到单一领域泛化模型中。为了进一步缓解类别不平衡,我们执行基于 GPT-5.5 的伪标签生成,并且我们最终的每个站点的合并模型在推理时不使用任何类别先验校正。所得模型在 MoCha 挑战赛的未见站点设置下进行评估,使用 Macro F1 作为主要评估指标。我们的方法在隐藏测试集上取得了 0.57 的宏 F1 分数,推理时仅使用 637K 个活跃参数,在 MoCha 2026 挑战赛的 58 个排行榜条目中排名第三。该挑战赛吸引了来自 112 名参与者的 1,669 份提交,并提供了由 Machine Medicine Technologies 赞助的奖金。
英文摘要
In this work, we introduce language-aligned motion representations for domain-generalizable UPDRS-Gait severity estimation, aiming to learn semantically structured motion features that generalize across heterogeneous clinical domains. We first learn motion representations using a Bi-GRU backbone that captures the temporal dynamics of SMPL sequences. Prior to model training, motion captions are generated offline using Qwen2.5-7B-Instruct. The backbone is then trained with both classification and text-alignment objectives to learn discriminative and semantically structured motion representations while accounting for the class imbalance present in the training data. We subsequently adapt the learned backbone independently to each source domain so that the model can capture domain-specific motion characteristics. The resulting source-specific models are then merged at the parameter level to consolidate complementary knowledge across source domains into a single domain-generalized model. To further mitigate class imbalance, we perform GPT-5.5-based pseudo labeling, and our final merged models for each site do not use any class-prior correction during inference. The resulting model is evaluated under the unseen-site setting of the MoCha Challenge, using Macro F1 as the primary evaluation metric. Our method achieves a macro-F1 of 0.57 on the hidden test set with only 637K active parameters at inference, ranking 3rd among 58 leaderboard entries in the MoCha 2026 Challenge. The challenge attracted 1,669 submissions from 112 participants and offered monetary prizes sponsored by Machine Medicine Technologies.
Comments3rd Place Solution to the MoCha 2026 Challenge at ECCV 2026