arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

对日常活动进行分类需要姿势,重建它们需要动作

Classifying daily activities needs posture, reconstructing them needs motion

Arefeh Farahmandi, Gunnar Blohm

arXiv 2607.13216首次发表:更新:

发表机构

Centre for Neuroscience Studies; Queen’s University(神经科学研究中心; 女王大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨人类如何快速分类日常活动动作,比较TMPs、勒让德多项式系数和自动编码器潜在嵌入三种策略,发现身体姿势和关键关节是分类判别特征,还揭示分类与重建动作时信息组织方式的分离及不同特征在其中的作用。

AI 中文摘要

人类能轻松识别动作,即便视觉输入嘈杂复杂。但刺激中的何种信息使人能快速分类动作?尚无框架系统比较不同动作分析策略。本文使用MoVi数据集中16种日常活动的视频,比较了三种策略:将动作分解为时间平滑基函数加权和的时间运动基元(TMPs)、将关节坐标轨迹投影到正交多项式基上的勒让德多项式系数、自动编码器潜在嵌入。勒让德系数和TMPs分类准确率最高,其次是自动编码器。发现了两个用于动作分类的判别特征,最具信息性的是身体的一般姿势,还识别出9个对动作分类最具预测性的关键关节。有趣的是,良好的分类准确率并不自动带来良好的动作生成:重建动作时,TMPs保留时间动态并产生自然运动,而勒让德系数重建仅保留平均姿势且显得僵硬。这些结果揭示了动作信息组织方式的分离:身体的静态配置足以分类执行的活动,但重建动作如何展开需要动作的时间动态。这种区别阐明了视觉系统可能依赖哪些特征进行快速动作识别,并表明姿势特征可在临床应用中实现高效动作筛选,而动态信息在以动作为目标的地方仍然至关重要。

英文摘要

Humans recognize movements effortlessly, even from noisy and complex visual input. But what information in the stimulus allows humans to rapidly classify movements? No framework has systematically compared different strategies of movement analysis to address this question. Here, we used videos of 16 daily activities from the MoVi dataset and compared three strategies: Temporal Movement Primitives (TMPs), which decompose movements into weighted sums of temporally smooth basis functions; Legendre polynomial coefficients, which project joint-coordinate trajectories onto an orthogonal polynomial basis; and Autoencoder latent embeddings. Legendre coefficients and TMPs achieved the highest classifier accuracy, followed by autoencoders. We found two discriminative features for movement classification. The most informative is the general posture of the body, the average spatial configuration that distinguishes one activity from another. Additionally, we identified 9 critical joints that are most predictive for movement classification. Interestingly, good classification accuracy did not automatically lead to good movement generation: when we reconstructed movements for each activity, TMPs preserved the temporal dynamics and produced perceptually natural motion, whereas reconstructions from Legendre coefficients retained only the average posture and appeared frozen. These results reveal a dissociation in how movement information is organized: the static configuration of the body suffices to classify what activity is performed, but the temporal dynamics of movement are required to reconstruct how it unfolds. This distinction clarifies which features the visual system may rely upon for rapid action recognition, and suggests that postural features could enable efficient movement screening in clinical applications, while dynamic information remain essential wherever movement generation is the goal.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑