arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19826cs.CV

MoAKE:通过动作知识专家混合实现统一的一体化动作质量评估

MoAKE: Toward Unified All-in-One Action Quality Assessment via Mixture of Action Knowledge Experts

  • College of Computer and Data Science, Fuzhou University(福州大学计算机与数据科学学院)
  • Fuzhou University(福州大学)
  • Engineering Research Center of Big Data Intelligence, Ministry of Education(教育部大数据智能工程研究中心)
  • School of Intelligence Science and Technology, University of Science and Technology Beijing(北京科技大学智能科学与技术学院)
  • University of Science and Technology Beijing(北京科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Huangbiao Xu, Huanqi Wu, Xiao Ke, Jiaxin Cai, Junyi Wu, Jinglin Xu

AI总结:

研究一体化动作质量评估难题,提出MoAKE框架,通过学习互补专家、定制原型及建模时间动态来减轻知识转移,在多数据集实验中显著优于现有方法,实现了零/少样本评估下的一致泛化。

AI中文摘要:

动作质量评估(AQA)旨在从动作视频中客观评估性能质量。大多数现有方法遵循“一对一”范式,为每种动作类型训练单独的模型。这种设置限制了实际部署,因为它需要先验动作类型知识来选择相应模型,并且在不同动作上泛化性差。为解决这些限制,我们研究了一体化AQA这一具有挑战性的任务,旨在在单个统一模型中评估异构动作。我们提出了一种新颖的动作知识专家混合(MoAKE)框架,以减轻动作间大语义差异导致的负面知识转移。MoAKE学习互补专家,在共享语义空间中捕获不同动作模式,并动态聚合其知识以适应输入动作评估。每个专家都用片段感知原型进行定制以处理不同时间长度,还配备了自适应段内和段间关系建模(AIISRM)模块来对多粒度时间动态进行建模。此外,我们为一体化以及零/少样本AQA建立了综合基准。在三个长期数据集上的广泛实验表明,MoAKE在一体化设置中显著优于现有方法,同时在零/少样本评估下在三个短期数据集上也实现了一致的泛化。代码可在该https网址获取。

英文摘要:

Action Quality Assessment (AQA) aims to objectively evaluate performance quality from action videos. Most existing methods follow a ``one-by-one'' paradigm, training a separate model for each action type. This setting limits real-world deployment, as it requires prior action-type knowledge to select the corresponding model and suffers from poor generalization across diverse actions. To address these limitations, we study the challenging task of all-in-one AQA, which aims to assess heterogeneous actions within a single unified model. We propose a novel Mixture of Action Knowledge Experts (MoAKE) framework, designed to mitigate negative knowledge transfer caused by large semantic discrepancies among actions. MoAKE learns complementary experts that capture diverse action patterns within a shared semantic space and dynamically aggregates their knowledge to adapt the assessment to the input action. Each expert is tailored with segment-aware prototypes to handle varying temporal lengths, together with an Adaptive Intra- and Inter-Segment Relationship Modeling (AIISRM) module to model multi-granularity temporal dynamics. Furthermore, we establish comprehensive benchmarks for all-in-one as well as zero/few-shot AQA. Extensive experiments on three long-term datasets demonstrate that MoAKE significantly outperforms existing methods in the all-in-one setting, while also achieving consistent generalization on three short-term datasets under zero/few-shot evaluation. Code is available at https://github.com/XuHuangbiao/MoAKE.

补充信息

↑