arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09169cs.CV

TSR-Ego:用于自中心3D人体姿态估计的时间引导立体细化框架

TSR-Ego: Temporally Guided Stereo Refinement Framework for Egocentric 3D Human Pose Estimation

Md Mushfiqur Azam, John Quarles, Kevin Desai

首次发表
浏览论文内容

中文总结 AI 辅助

针对自中心3D人体姿态估计难题,提出TSR-Ego框架,结合短期运动证据与投影引导特征采样,通过因果深度可分离时间卷积和多种自注意力机制细化3D关节查询,实验证明其性能领先,尤其在真实世界序列中优势明显。

中文摘要 AI 辅助

从头戴式立体相机进行自中心3D人体姿态估计具有挑战性,因为存在鱼眼失真、严重自遮挡和相机视野外身体关节的频繁截断。近期方法通过热图提升、立体对应和基于Transformer的细化提高了性能,但常依赖帧局部证据或仅将时间信息用作辅助姿态级上下文。我们提出TSR-Ego,一个将短期运动证据与投影引导特征采样相结合的时间引导立体框架。该模型首先使用因果深度可分离时间卷积丰富密集立体特征图,然后通过时间自注意力、关节自注意力和鱼眼可变形立体交叉注意力来细化学习到的3D关节查询。实验表明TSR-Ego在UnrealEgo2和UnrealEgo-RW上展现了领先性能,在真实世界序列上有显著提升。

英文摘要

Egocentric 3D human pose estimation from head-mounted stereo cameras is challenging due to fisheye distortion, severe self-occlusion, and frequent truncation of body joints outside the camera field of view. Recent stereo egocentric methods have improved performance through heatmap lifting, stereo correspondence, and transformer-based refinement, but they often rely heavily on frame-local evidence or use temporal information only as auxiliary pose-level context. This limits robustness when current-frame stereo cues are weak, occluded, or ambiguous. We propose TSR-Ego, a temporally guided stereo framework that couples short-term motion evidence with projection-guided feature sampling. The model first enriches dense stereo feature maps using a causal depthwise-separable temporal convolution, allowing past visual evidence to influence the feature space before deformable cross-attention. A single-stage causal stereo decoder then refines learned 3D joint queries through temporal self-attention, joint self-attention, and fisheye deformable stereo cross-attention, using the evolving pose estimate to generate 2D sampling references. Unlike methods that apply temporal reasoning mainly after pose prediction, TSR-Ego uses motion context to shape both the sampled stereo features and the joint representations while preserving online inference without future frames. Experiments on UnrealEgo2 and UnrealEgo-RW show state-of-the-art performance, with especially strong gains on real-world sequences.

发表机构

  • The University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)

机构由 AI 辅助整理,请以论文原文为准。

↑