发表机构
National Institute of Informatics(国立信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对MLLMs长视频理解的时序组织缺失问题,提出无训练的T³框架,通过分层时序树与“回答-检索-探索”循环实现粗到细的线索搜索,在三个基准数据集上提升了Qwen2.5-VL-7B的性能。
AI 中文摘要
多模态大语言模型(MLLMs)因上下文长度有限,长视频理解仍具挑战性。均匀采样可能遗漏关键时刻,而基于智能体的视频帧理解方法常独立评估各帧,忽略视频的时序组织。理想情况下,证据选择应模仿人类回答长视频问题的方式:先从全局语境中定位相关片段,再放大至局部对象与细节。我们提出Temporal Tree of Thought(T³),一种无需训练的自适应粗到细长视频理解框架。T³通过递归时序约束聚类构建与问题无关的分层时序树,每个节点代表含信息关键帧的连续片段。推理阶段,T³执行“回答-检索-探索”循环:对粗粒度代表性帧推理,证据不足时生成搜索语句,并扩展相关分支以获取更细粒度证据。该过程自适应将搜索目标从时序区域转向特定对象与视觉细节,助力视频理解。在VideoMME、LongVideoBench和LVBench上的实验表明,在相同帧预算下,T³分别将Qwen2.5-VL-7B的性能提升0.5%、4.6%和4.4%,验证了结构化时序推理的有效性。
英文摘要
Long-video understanding remains challenging for Multimodal Large Language Models (MLLMs) due to limited context length. Uniform sampling may miss crucial moments, while agent-based frame video understanding methods often evaluate frames independently, overlooking the temporal organization of videos. Ideally, evidence selection should mimic how humans answer questions about long videos: first locating the relevant segment from the global context, then zooming into local objects and details. We propose Temporal Tree of Thought T^3, a training-free framework for adaptive coarse-to-fine long-video understanding. T^3 constructs a question-agnostic hierarchical temporal tree via recursive temporally constrained clustering, where each node represents a contiguous segment with an informative key frame. During inference, T^3 performs an answer-retrieve-explore loop: it reasons over coarse representative frames, generates a search statement when evidence is insufficient, and expands relevant branches for finer-grained evidence. This process adaptively shifts the search target from temporal regions to specific objects and visual details to help video understanding. Experiments on VideoMME, LongVideoBench, and LVBench show that T^3 improves Qwen2.5-VL-7B by 0.5%, 4.6%, and 4.4%, respectively, under the same frame budget, demonstrating the effectiveness of structured temporal reasoning.
CommentsAccept by EMNLP2026