arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentSTAR:从单目视频进行智能体式形状重建与跟踪

AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos

Kirill Mazur, Nikita Karaev, Matthew Chang, Jitendra Malik, Nur Muhammad "Mahi'' Shafiullah

arXiv 2609.24487首次发表:更新:

发表机构

Amazon FAR (Frontier AI and Robotics)(亚马逊前沿人工智能与机器人部门)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出AgentSTAR方法,利用视觉-语言模型智能体通过分析-合成循环从单目视频重建和跟踪物体形状与位姿,在ARCTIC和HOT3D数据集上超越现有基线。

AI 中文摘要

在这项工作中,我们提出了一种通过智能体式分析-合成方法从视频中进行形状重建与跟踪的方法。与先前先估计密集像素对应关系、再从中恢复物体运动的方法不同,我们的方法推断出一个结构化的3D物体模型,包括其几何和运动学结构,并利用该模型随时间优化物体跟踪估计。在我们的优化循环中,一个视觉-语言模型(VLM)智能体通过渲染-比较循环迭代地细化形状或广义位姿,将粗略的视觉推理与数值位姿优化相结合,以实现精确的状态估计。这种结构化公式使我们的方法能够在无需依赖像素匹配目标的情况下,跟踪大幅运动、关节运动和严重遮挡。在定量评估中,在ARCTIC数据集上,我们的方法在关节物体的3D点跟踪方面大幅优于最先进的基线方法;在HOT3D数据集上,我们的方法优于所有评估的刚体跟踪基线方法。

英文摘要

In this work, we present a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis. Unlike prior methods which first estimate dense pixel correspondences and then recover object motion from them, our method infers a structured 3D object model, including its geometry and kinematic structure, and uses this model to optimise object track estimates over time. In our optimisation loop, a Vision-Language Model (VLM) agent iteratively refines shape or generalised pose through a render-and-compare loop, combining coarse visual reasoning with numerical pose optimisation for precise state estimation. This structured formulation enables our method to track through large motion, articulation, and severe occlusion without relying on pixel-matching objectives. Quantitatively, on ARCTIC, our method substantially outperforms state-of-the-art 3D point-tracking baselines for articulated objects, and on HOT3D it outperforms all evaluated rigid-object tracking baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑