arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高效跟踪与理解对象变换

Efficient Tracking and Understanding Object Transformations

Yihong Sun, Bharath Hariharan

arXiv 2607.19743首次发表:更新:

发表机构

Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何高效跟踪与理解对象变换,提出FluxGraph方法,利用SAM2内部多掩码差异触发变换检测,无需跟踪视频所有实体,相比TubeletGraph速度大幅提升,在多个数据集上也有显著加速且性能良好。

AI 中文摘要

通过状态变换跟踪对象对于理解现实世界动态至关重要。然而,现有方法计算成本高昂。TubeletGraph虽有强大能力,但推理成本高(如在VOST上每个对象帧约4.4秒)无法实时部署。我们发现其开销源于构建输入视频的时空分区:对每帧密集计算实体分割,且跟踪场景中每个实体,成本随场景复杂度而非感兴趣的变换数量增加。为解决这些,我们提出FluxGraph,它利用SAM2的内部多掩码差异作为变换检测的轻量级触发器,无需跟踪给定视频中的所有实体。在VOST上,FluxGraph比TubeletGraph快约3.3倍,同时提升跟踪性能并保持状态图质量。在VSCOS、M³-VOS和DAVIS17上也有3.7 - 10.7倍的加速且性能良好。代码可公开获取。

英文摘要

Tracking objects through state transformations is essential for understanding real-world dynamics. However, existing methods are computationally expensive. TubeletGraph recently showed impressive capabilities, but its inference cost (~$4.4$ seconds per object-frame on VOST) precludes any real-time deployment possibilities. We observe that TubeletGraph's overhead arises from building a spatiotemporal partition of the input video: (1) entity segmentation is computed densely for every frame regardless of whether a transformation occurs, and (2) every entity in the scene is tracked, scaling cost with scene complexity rather than the number of transformations of interest. To address both, we propose FluxGraph, a reactive variant that uses SAM2's internal multi-mask disagreement as a lightweight trigger for transformation detection, and removes the need for tracking all entities in the given video. FluxGraph is ~$3.3\times$ faster than TubeletGraph on VOST while improving tracking performance and preserving state graph quality. Furthermore, we also observe consistent speedups of $3.7-10.7\times$ across VSCOS, M$^3$-VOS, and DAVIS17 while maintaining performance. Code is publicly available at https://github.com/YihongSun/FluxGraph.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑