FineMoLA:基于片段级监督的细粒度动作-语言对齐研究
FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision
- Purdue University(普渡大学)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有动作-语言数据集仅提供片段级监督导致细粒度对齐不足的问题,提出弱监督框架FineMoLA,将动作-语言对齐建模为最优传输问题,在SnapMoGen上的实验显示其动作-文本定位性能优于基线。
AI中文摘要:
文本条件下的人体动作生成随着大规模动作-语言数据集的出现取得了快速进展。然而,即使是具有丰富长文本描述的数据集,通常也仅提供片段级监督,未明确动作帧与语言之间的时间对应关系,这限制了细粒度动作-文本定位和时间精确生成。我们提出FineMoLA,这是一种弱监督框架,可直接从片段级注释学习细粒度帧-短语对应关系。我们的方法首先将长文本描述分割为包含动作的短语,然后将动作-语言对齐建模为最优传输问题,该问题在全局约束下自然建模动作帧与文本之间的多对多关系。结合熵正则化和Sinkhorn迭代,FineMoLA可高效推断伪帧级对齐,无需人工标注。在SnapMoGen数据集上的实验表明,学习到的对齐在动作-文本定位方面优于基线方法。
英文摘要:
Text-conditioned human motion generation has made rapid progress with the emergence of large-scale motion--language datasets. However, even datasets with rich long-form descriptions typically provide supervision only at the clip level, without explicit temporal correspondence between motion frames and language. This limits fine-grained motion--text grounding and temporally precise generation. We propose FineMoLA, a weakly supervised framework that learns fine-grained frame--phrase correspondence directly from clip-level annotations. Our method first segments long-form descriptions into action-bearing phrases, and then formulates motion--language alignment as an optimal transport problem, which naturally models many-to-many relations between motion frames and text under global constraints. With entropic regularization and Sinkhorn iterations, FineMoLA efficiently infers pseudo frame-level alignments without human labeling. Experiments on SnapMoGen demonstrate that the learned alignments outperform baselines in motion--text grounding.