VARPose:基于视觉自回归建模的灵活2D姿态密集化以提升3D姿态提升性能
VARPose: Flexible 2D Pose Densification via Visual Autoregressive Modeling for Enhanced 3D Lifting
- School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院)
- The Technology Innovation Center for Collaborative Applications of Natural Resources Data in GBA, MNR(自然资源部粤港澳大湾区自然资源数据协同应用技术创新中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
VARPose基于VAR建模,通过GPT分词器和UniSkelar自回归模型实现2D姿态灵活密集化,在3D姿态估计等下游任务上性能优于现有方法且可泛化到新粒度。
AI中文摘要:
视觉自回归建模(VAR)已通过下尺度预测在自然图像生成领域表现出色,但其在人体骨骼这类拓扑结构化数据上的应用仍未被探索。本文提出VARPose以自适应地密集化2D稀疏姿态,从而为3D姿态提升模型补充可用的解剖信息。核心贡献有两点:第一,我们引入了粒度无关姿态分词器(GPT),其采用单一混合码本和残差量化策略,将不同密度的姿态编码为统一的多尺度离散表示,结果表明该表示具有强泛化性;通过将表示与投影解耦,我们可利用冻结码本和重新训练的解码器成功解码新的姿态粒度。第二,我们提出UniSkelar,这是一种统一自回归模型,将“关节密度”视为“尺度”,以最稀疏的姿态为条件,按由粗到细的方式学习预测下一个密度级别的标记序列。VARPose不仅优于现有最先进方法且能泛化到未见过的粒度,还通过2D姿态密集化为下游任务(如3D姿态估计和人体网格恢复)带来切实的性能提升,代码和模型可在指定网址获取。
英文摘要:
Visual AutoRegressive Modeling (VAR) has excelled in natural image generation via next-scale prediction, but its use on topology-structured data like human skeletons is still unexplored. VARPose is proposed to adaptively densify 2D sparse poses, thereby enriching the anatomical information available for 3D lifting models. Our core contributions are twofold. First, we introduce a Granularity-agnostic Pose Tokenizer (GPT), which employs a single hybrid codebook and a residual quantization strategy to encode poses of varying densities into a unified, multi-scale discrete representation. Our results demonstrate the strong generalizability of this representation. By decoupling the representation from the projection, we can successfully decode novel pose granularities using a frozen codebook with a retrained decoder. Second, we propose UniSkelar, a unified autoregressive model that treats "joint density" as "scale". UniSkelar learns to predict the token sequence for the next density level in a coarse-to-fine manner, conditioned on the sparsest pose. VARPose not only outperforms state-of-the-art methods and generalizes to unseen granularities, but also confers tangible performance gains on downstream tasks, such as 3D Pose Estimation and Human Mesh Recovery, through 2D pose densification. Our code and model are available at https://github.com/BRL-SYSU/VARPose.git.