arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32462cs.CV

运动-语言模型能否理解结构?STRIDE:评估评估者

Can Motion-Language Models Ground Structure? STRIDE for Evaluating the Evaluators

Lixing Tan, Qing Xia, Yuting Guo, Shuai Li, Aimin Hao

首次发表
浏览论文内容

中文总结 AI 辅助

针对运动-语言模型评估器对语言结构理解不足的问题,提出STRIDE基准系统评估时间顺序、镜像反射和动作身份,发现现有评估器存在严重缺陷,并通过结构硬负样本改进对比学习显著提升性能。

中文摘要 AI 辅助

运动-语言模型通常由运动-语言评估器进行评分,但这些评估器对语言结构的理解程度仍不明确。为此,我们引入了基于时间顺序、镜像反射和动作身份的结构接地诊断评估(STRIDE)基准,以系统评估评估器追踪时间顺序、镜像反射和动作身份的能力。STRIDE包含5,869个三元组,每个三元组由一个运动、其原始描述和一个扰动描述组成,涵盖短描述和长描述。我们对描述对进行似然平衡以减少纯文本偏差,并在无关运动下估计每个评估器的描述偏好,以衡量匹配运动相对于基线的判别增益。我们的实验揭示了被审计评估器存在弱结构接地和严重的镜像敏感性缺陷,而常用的评估协议未能暴露这些问题。为理解这些局限为何在标准测试中未被发现,我们更仔细地检查了评估器。我们发现,仅凭文本先验即可解决现有数据集上的朴素扰动测试,而常见的检索和分布度量几乎对镜像真实运动引入的结构破坏无响应。这些发现提示了一种自然的干预措施:结构硬负样本。我们的实验表明,对对比学习进行简单修改即可显著提升时间顺序和镜像反射的性能。该基准和代码将公开发布。

英文摘要

Motion-language models are typically scored by motion-language evaluators, but how well these evaluators ground language structure remains unclear. Here, we introduce the Structure grounding via Temporal-order, Reflection, and Identity Diagnostic Evaluation (STRIDE) benchmark to systematically evaluate the ability of evaluators to track temporal order, mirror reflection, and action identity. STRIDE comprises $5{,}869$ triples, each consisting of a motion, its original caption, and a perturbed caption, spanning both short and long descriptions. We likelihood-balance caption pairs to reduce text-only bias and estimate each evaluator's caption preference under unrelated motions to measure the discrimination gain from matched motions relative to this baseline. Our experiments reveal weak structural grounding and severe deficits in mirror sensitivity among the audited evaluators, which commonly used evaluation protocols fail to expose. To understand why these limitations go undetected in standard tests, we examine the evaluators more closely. We find that text-only priors alone can solve naive perturbation tests on existing datasets, while common retrieval and distributional metrics barely respond to structural corruption introduced by mirroring ground-truth motions. These findings suggest a natural intervention: structural hard negatives. Our experiments show that a simple modification to contrastive learning substantially improves performance on temporal order and mirror reflection. The benchmark and code will be released.

↑