模型作为工具:用于统一多模态视觉跟踪的智能体协调框架
Models as Tools: An Agentic Coordination Framework for Unified Multimodal Visual Tracking
浏览论文内容
中文总结 AI 辅助
针对现有视觉跟踪方法的瓶颈,提出ACTrack智能体协调框架,整合异构模型工具的互补优势,仅用30%可训练参数便在多类基准上实现性能超越。
中文摘要 AI 辅助
当前大多数视觉跟踪器采用仅在跟踪数据集上训练的基于匹配的架构,其性能提升高度依赖输入上下文的长度,现已达到瓶颈。尽管高性能跟踪越来越依赖基础模型,但现有方法将其整体使用,要么将基础模型适配为跟踪器,要么将分割基础模型修改为跟踪流程,未能利用互补优势。基于匹配的跟踪器擅长实例级对应,但缺乏语义判别和细粒度前景感知;而分割基础模型能生成精确掩码,但在实例判别和多模态扩展方面存在不足。两种范式均缺乏长期跟踪的纠错能力。为解决这些问题,我们提出ACTrack,一种将异构模型视为可调用工具的智能体协调框架,采用事件触发机制。ACTrack协调四种工具:用于目标判别的基于跟踪器的实例匹配工具、用于掩码衍生运动先验的SAM3运动工具、用于检测干扰项和实例冲突线索的SAM3感知工具,以及仅在持续冲突时激活以缓解误差累积的VLM重提示工具。我们设计了完整的工具调用触发机制和工具间协调机制,以充分整合不同模型工具的互补优势。实验表明,ACTrack在8个RGB基准上显著超越最强、最大的跟踪器。此外,参数高效的适配策略实现了工具间的参数共享与复用,仅用30%的可训练参数即可实现统一多模态跟踪,并在LasHeR、VisEvent、TNL2K和DepthTrack等多模态基准上显著优于现有方法。
英文摘要
Most current visual trackers adopt a matching-based architecture trained exclusively on tracking datasets, whose performance gains depend heavily on the length of the input context, and have now reached a bottleneck. While high-performance tracking increasingly relies on foundation models, existing methods use them monolithically, adapting a foundation model into a tracker or modify a segmentation foundation model into a tracking pipeline, which fails to exploit complementary strengths. Matching-based trackers excel at instance-level correspondence but lack semantic discrimination and fine-grained foreground perception, whereas segmentation foundation models produce precise masks yet struggle with instance discrimination and multimodal extension. Both paradigms also lack error-correction capabilities for long-term tracking. To address these issues, we propose ACTrack, an agentic coordination framework that treats heterogeneous models as invocable tools under an event-triggered mechanism. ACTrack coordinates a Tracker-based Instance Matching Tool for target discrimination, a SAM3 Motion Tool for mask-derived motion priors, a SAM3 Perception Tool for detecting distractors and instance-conflict cues, and a VLM Reprompt Tool activated only under persistent conflict to mitigate error accumulation. We design a complete tool-invocation trigger mechanism and an inter-tool coordination mechanism, enabling the complementary strengths of different model tools to be fully integrated. Experiments show that ACTrack substantially surpasses the strongest and the largest trackers on eight RGB benchmarks. Furthermore, a parameter-efficient adaptation strategy enables parameter sharing and reuse across tools, achieving unified multimodal tracking with only 30\% trainable parameters while substantially outperforming prior methods on multimodal benchmarks such as LasHeR, VisEvent, TNL2K, and DepthTrack.
发表机构
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。