发表机构
University of North Carolina at Chapel Hill; Sony(北卡罗来纳大学教堂山分校; 索尼)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究有基础的长视频问答问题,提出VideoTreeSearch框架,通过构建自适应时间树及可学习的离散操作让智能体导航,经监督微调与强化学习训练,在多个基准测试中表现出色,证明自校正分层搜索是关键机制。
AI 中文摘要
有基础的长视频问答(Grounded LVQA)需要在定位支持答案的短证据区间的同时回答关于长视频的问题。最近的智能体方法将此任务框架化为单作物视频(开始,结束)动作的多轮探索,支持从粗到细的缩小,但没有从细到粗回溯的原语。因此,这些智能体通常过早收敛,无法从早期错误中恢复。我们提出了VideoTreeSearch(VTS),一个将有基础的LVQA框架化为对自适应时间树的迭代自校正搜索的框架。VTS从视觉场景边界构建一个非均匀树,使每个节点对应一个语义连贯的片段,并训练一个智能体通过四种离散操作在树中导航:放大、缩小、移动和回答。这些操作将回溯和恢复暴露为显式的、可学习的原语,而不是隐式行为。为了训练这种导航,我们引入了一个轨迹合成管道,该管道生成通过树的多步路径,包括故意绕入错误分支然后恢复。我们使用这些轨迹进行监督微调,然后使用基础和答案准确性奖励进行强化学习。在三个有基础的LVQA基准(CG-Bench、Haystack-LVBench、Haystack-Ego4D)上,VTS在CG-Bench上比最强的先前智能体方法高出+12.5 mIoU,在Haystack-Ego4D上高出+7.4 T-F1。学习到的策略也转移到了一般的长视频问答中,在Video-MME、MLVU和LVBench上超过了所有先前的智能体基线,提高了多达+7.1个准确性点。消融实验证实,自校正分层搜索是这些收益背后的核心机制:去除自适应下降或显式回溯会显著降低性能。代码可在此https URL获取。
英文摘要
Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while localizing the short evidence interval that supports the answer. Recent agentic methods frame this task as multi-turn exploration with a single crop_video(start, end) action, which supports coarse-to-fine narrowing but provides no primitive for fine-to-coarse backtracking. As a result, these agents typically converge prematurely and cannot recover from an early mistake. We propose VideoTreeSearch (VTS), a framework that casts grounded LVQA as iterative self-correcting search over an adaptive temporal tree. VTS constructs a non-uniform tree from visual scene boundaries so that each node corresponds to a semantically coherent segment, and trains an agent to navigate the tree through four discrete operations: zoom_in, zoom_out, shift, and answer. These operations expose backtracking and recovery as explicit, learnable primitives rather than implicit behaviors. To train this navigation, we introduce a trajectory synthesis pipeline that produces multi-step paths through the tree, including deliberate detours into incorrect branches followed by recovery. We use these trajectories for supervised fine-tuning, followed by reinforcement learning with grounding and answer-accuracy rewards. On three Grounded LVQA benchmarks (CG-Bench, Haystack-LVBench, Haystack-Ego4D), VTS outperforms the strongest prior agentic methods by +12.5 mIoU on CG-Bench and +7.4 T-F1 on Haystack-Ego4D. The learned policy also transfers to general long-video QA, surpassing all prior agentic baselines on Video-MME, MLVU, and LVBench by up to +7.1 accuracy points. Ablations confirm that self-correcting hierarchical search is the central mechanism behind these gains: removing either adaptive descent or explicit backtracking substantially degrades performance. Code is available at https://github.com/CeeZh/VTS.