arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

零样本基于骨架的动作预测

Zero-Shot Skeleton-Based Action Anticipation

Hongsong Wang, Pengbo Yan, Yang Zhang, Qiuxia Lai

arXiv 2608.14243首次发表:更新:

发表机构

School of Computer Science and Engineering, Southeast University; Southeast University - Monash University Joint Graduate School; School of Computer Science and Software Engineering, Shenzhen University; State Key Laboratory of Media Convergence and Communication, Communication University of China(东南大学计算机科学与工程学院; 东南大学-莫纳什大学联合研究生院; 深圳大学计算机与软件学院; 中国传媒大学媒体融合与传播国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对现有基于骨架的动作预测无法泛化到新动作的问题,提出零样本基于骨架的动作预测新任务,构建基线模型与基准协议,在NTU RGB+D上实现高零样本准确率,为相关真实系统研究奠定基础。

AI 中文摘要

动作预测(Action Anticipation, AA)旨在通过部分观测识别正在进行的人类或类人机器人动作,使机器人能在动作完成前预测意图。尽管基于骨架的动作预测具有效率优势,但现有方法假设训练阶段已见过所有动作类别,这限制了其在必然出现新动作类别的真实场景中的部署。为解决这一差距,本文研究零样本基于骨架的动作预测(Zero-Shot Skeleton-Based Action Anticipation, ZS-SkAA)这一新任务,该任务要求仅用有限的早期骨架序列识别未见动作类别,融合了部分观测、时间动态和零样本泛化的挑战。为为ZS-SkAA奠定基础研究,本文引入:(1)一个基线模型,包含时空特征提取器和互信息估计与最大化模块,该基线模型通过估计并最大化部分视觉特征与语义类嵌入跨模态的互信息,明确对齐两者,提升对未见类别的泛化能力;(2)一个使用NTU RGB+D数据集的基准协议,该协议经调整后可用于严格的ZS-SkAA评估。实验表明,本文提出的模型作为ZS-SkAA的强基线具有有效性,在NTU RGB+D上实现了高零样本准确率。本研究将ZS-SkAA确立为需要对新动作进行泛化的真实系统的重要研究方向。

英文摘要

Action anticipation (AA) aims to recognize ongoing human or humanoids actions from partial observations, enabling robots to predict intentions before the actions are completed. Although skeleton-based AA offers efficiency advantages, existing approaches assume that all action classes are seen during training, which limits their deployment in real-world scenarios where novel actions inevitably arise. To address this gap, we study the new task of Zero-Shot Skeleton-Based Action Anticipation (ZS-SkAA). This task requires recognizing unseen action classes using only limited early-stage skeleton sequences, combining the challenges of partial observations, temporal dynamics, and zero-shot generalization. To establish foundational research for ZS-SkAA, we introduce:(1) A baseline model comprising a spatio-temporal feature extractor and a mutual information estimation and maximization module. This baseline model explicitly aligns partial visual features with semantic class embeddings across modalities by estimating and maximizing their mutual information, enhancing generalization to unseen classes.(2) A benchmark protocol using the NTU RGB+D dataset, which is adapted for rigorous ZS-SkAA evaluation. Experiments demonstrate the effectiveness of our model as a strong baseline for ZS-SkAA, achieving high zero-shot accuracy on NTU RGB+D. This work establishes ZS-SkAA as a vital research direction for real-world systems requiring generalization to novel actions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑