arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12187cs.CVcs.AI

HSTGFormer:用于3D人体姿态估计的超时空图变换器

HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation

  • Durham University(杜伦大学)

机构由 AI 辅助整理,请以论文原文为准。

Ruochen Li, Shuang Chen, Wenke E, Farshad Arvin, Amir Atapour-Abarghouei

中文总结 AI 辅助

本文提出HSTGFormer框架,通过超时空图与自适应双尺度时间图实现时空耦合推理,在Human3.6M等数据集上以高准确率与计算效率完成3D人体姿态估计。

中文摘要 AI 辅助

基于Transformer的方法在单目3D人体姿态估计中已取得优异性能,但现有多数方法将空间与时间推理组织为独立阶段,可能削弱人体运动固有的统一时空相互依赖性,并在时间建模前压缩帧级结构信息。本文提出HSTGFormer,一种增强图的Transformer框架,将时空推理重新表述为对关节-时间节点的局部耦合图聚合。具体而言,HSTGFormer引入超时空图(Hyper Spatial-Temporal Graph, HSTG),通过将逐帧骨架图扩展至时间邻域,将全局时空推理分解为围绕单个关节-时间节点的局部时空感受野,从而在保留局部结构运动信息的同时实现感知结构的耦合推理。该模型还融入自适应双尺度时间图(Adaptive Dual-Scale Temporal Graph, ADSTG),以互补的短程与长程窗口捕获关节特异性时间依赖关系。轻量逐节点融合模块进一步为每个关节-时间节点自适应整合两种图表示。在Human3.6M与MPI-INF-3DHP数据集上的实验表明,HSTGFormer在实现高准确率的同时具备高计算效率。

英文摘要

Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spatial and temporal reasoning as separate stages, which may weaken unified spatial-temporal interdependencies inherent in human motion and compress frame-level structural information before temporal modelling. In this paper, we propose HSTGFormer, a graph-enhanced Transformer framework that reformulates spatial-temporal reasoning as localised coupled graph aggregation over joint-time nodes. Specifically, HSTGFormer introduces a Hyper Spatial-Temporal Graph (HSTG), which decomposes global spatial-temporal reasoning into local spatial-temporal receptive fields around individual joint-time nodes by extending per-frame skeleton graphs into temporal neighbourhoods, thereby enabling structure-aware coupled reasoning while preserving local structural motion information. It further incorporates an Adaptive Dual-Scale Temporal Graph (ADSTG) to capture joint-specific temporal dependencies over complementary short- and long-range windows. A lightweight node-wise fusion module further adaptively integrates the two graph representations for each joint-time node. Experiments on Human3.6M and MPI-INF-3DHP show that HSTGFormer achieves strong accuracy with high computational efficiency.

补充信息

↑