发表机构
University of Bath; University of Washington(巴斯大学; 华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对交通异常理解需求,提出TAU-Agent智能体式检索增强框架,调度两类视觉工具获取证据,结合微调视觉语言模型推理,在AI City 2026挑战赛多赛道取得对应排名成绩。
AI 中文摘要
交通异常理解(TAU)要求模型与系统能够检测、推理并解释交通视频中的异常事件。为应对这一挑战,我们提出TAU-Agent,这是一种用于交通异常理解的智能体式检索增强框架。给定任务查询后,一个核心检索智能体将调度两个视觉感知工具,即视频字幕工具与开放词汇跟踪工具,以检索并选择与查询相关的证据,包括字幕、时间区间及目标轨迹。所选证据、采样的视频帧以及输入查询,将被提供给经过监督微调的视觉语言模型,用于最终推理与答案生成。我们在2026年AI City挑战赛的域内与域外基准上对TAU-Agent进行评估,TAU-Agent在Track 3取得0.6779分、Track 7取得0.3998分、Track 8取得67.9275分,分别排名第二、第十二与第五。代码可在此URL获取。
英文摘要
Traffic Anomaly Understanding (TAU) requires models and systems to detect, reason about, and explain anomalous events in transportation videos. To address this challenge, we propose TAU-Agent, an agentic retrieval-augmented framework for traffic anomaly understanding. Given a task query, a central retrieval agent orchestrates two visual perception tools, namely a Video Captioning Tool and an Open-Vocabulary Tracking Tool, to retrieve and select query-relevant evidence, including captions, temporal intervals, and object trajectories. The selected evidence, together with sampled video frames and the input query, is provided to a supervised fine-tuned vision-language model for final reasoning and answer generation. We evaluate TAU-Agent on both the in-domain and the out-of-domain benchmarks from the AI City Challenge 2026. TAU-Agent achieves scores of 0.6779 on Track 3, 0.3998 on Track 7, and 67.9275 on Track 8, ranking second, twelfth, and fifth, respectively. Code is available at: https://github.com/siri-rouser/TAU-Agent.