发表机构
Harvard University; NVIDIA; Nanyang Technological University(哈佛大学; 英伟达; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对多人场景下人类网格恢复难题,提出DETRAM统一框架,利用单个变压器解码器及可学习查询嵌入,能自动和依用户提示检测、重建与跟踪人类,在多数据集上取得领先跟踪结果及有竞争力的重建精度,实现端到端可训练的用户导向人体分析。
AI 中文摘要
在人类网格恢复(HMR)任务中,多人场景因实体众多且随时间出现遮挡而难以处理,尤其对于视频输入,需要可靠且一致地跟踪每个实体。现有方法依赖预训练的人体检测模块,增加了运行时间并限制了跟踪实体数量。我们提出了DETRAM,这是一个用于多人HMR和跟踪的统一框架,能跨时间自动并通过用户提示同时检测、重建和跟踪人类。DETRAM使用单个具有跨帧持久的身份一致的可学习查询嵌入的变压器解码器:检测查询发现新人,跟踪查询维护现有个体的姿态和形状,提示查询跟踪用户指定的身份。我们的方法在PoseTrack21、3DPW、BEDLAM和MuPoTS - 3D上取得了领先的跟踪结果,在BEDLAM和3DPW上具有有竞争力的重建精度,同时独特地支持多人场景中基于提示的个体跟踪。据我们所知,这是第一种在端到端可训练框架中将可提示性、多人HMR与跟踪统一起来的方法,实现了视频中用户导向的人体分析。
英文摘要
In the task of human mesh recovery (HMR), multi-person scenes are particularly difficult to handle due to the many entities that appear and occlusions between them over time. In particular for video inputs, there is a need to track each entity reliably and consistently. Existing methods rely on pretrained human detection modules, increasing their runtime and limiting the number of tracked entities. We present DETRAM, a unified framework for multi-person HMR and tracking that simultaneously detects, reconstructs, and tracks humans across time, both automatically and via user prompts. DETRAM uses a single transformer decoder with an identity-consistent set of learnable query embeddings that persist across frames: detection queries discover new people, tracking queries maintain pose and shape for existing individuals, and prompt queries follow user-specified identities. Our approach achieves state-of-the-art tracking results on PoseTrack21, 3DPW, BEDLAM, and MuPoTS-3D, and competitive reconstruction accuracy on BEDLAM and 3DPW, while uniquely supporting prompt-based tracking of individuals in multi-person scenes. To our knowledge, this is the first method to unify promptability and multi-person HMR with tracking in an end-to-end trainable framework, enabling user-directed human analysis in videos.