arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MRBench:用于人体动作-文本检索的综合基准

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval

Fulong Liu, Liang Xu, Chengqun Yang, Yuhao Zhang, Yichao Yan, Xiaokang Yang

arXiv 2608.07993首次发表:更新:

AI 中文总结

本文提出综合动作-文本检索基准MRBench,针对现有基准缺陷构建,提出轻量级粒度感知模型,提升混合粒度检索效果,为动作-语言对齐评估提供测试平台。

AI 中文摘要

人体动作-文本检索为评估跨模态对齐提供了严格手段。现有主流基准多以同质室内动作为主,存在动作分布失衡、文本过于简化且重复的问题,阻碍了跨领域和跨粒度对齐的可靠测量。为此,本文提出MRBench,这是一个综合动作-文本检索基准,具有异质动作、广泛且均衡的类别覆盖,以及可靠、有区分度的多粒度描述。MRBench通过精心设计的多阶段数据整理流程构建,该流程筛选并平衡候选数据、验证无歧义的语义对齐,并生成基于动作的多粒度描述。最终的基准包含3390个动作,来自动作捕捉、野外视频、合成视频和动作生成模型,覆盖118个细粒度类别。每个动作都配有简洁、标准且细粒度的描述,共产生10170个文本描述。对MRBench上代表性检索基线的广泛评估显示,存在显著的跨数据集泛化差距,且对查询粒度有明显的敏感性。本文提出一种轻量级粒度感知模型,以冻结的标准描述对齐检索模型为锚点。基于大语言模型(LLM)的简洁和细粒度文本描述为额外分支的粒度特定动作提取器和文本适配器提供伪监督。推理时,粒度感知分数融合整合全局和适配后的相似度,同时严格保持所有描述级别间的分数可比性。该模型在不损害标准描述性能的情况下,提升了混合粒度检索效果。本文认为MRBench为推进动作-语言对齐评估提供了综合测试平台。

英文摘要

Human motion-text retrieval provides a rigorous means of assessing cross-modal alignment. Prevailing benchmarks are dominated by homogeneous indoor motions, imbalanced motion distributions, and oversimplified, repetitive texts, which hinder the reliable measurement of cross-domain and cross-granularity alignment. We thus introduce MRBench, a comprehensive motion-text retrieval benchmark featuring heterogeneous motions, broad and balanced category coverage, and reliable, discriminative, multi-granular descriptions. MRBench is constructed through a meticulously designed multi-stage data curation pipeline, which filters and balances candidates, verifies unambiguous semantic alignment, and generates motion-grounded descriptions at multiple granularities. The resulting benchmark contains 3,390 motions drawn from motion capture, in-the-wild videos, synthetic videos, and motion generative models, covering 118 fine-grained categories. Each motion is paired with concise, standard, and fine-grained descriptions, yielding 10,170 captions. Extensive evaluations of representative retrieval baselines on MRBench reveal a substantial cross-dataset generalization gap and pronounced sensitivity to query granularity. We propose a lightweight granularity-aware model anchored at a frozen standard-caption-aligned retrieval model. LLM-based concise and fine-grained captions provide pseudo-supervision for extra-branch granularity-specific motion extractors and text adapters. For inference, granularity-aware score fusion integrates global and adapted similarities while strictly maintaining score comparability across all description levels. The resulting model improves mixed-granularity retrieval without compromising standard-caption performance. We believe that our MRBench provides a comprehensive testbed for advancing motion-language alignment evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑