发表机构
University of Science, VNU-HCM(胡志明市越南国立大学科学大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对当代视觉语言模型理解足球视频的挑战,提出TreeSoc框架,通过动态深度优先搜索机制分解查询,支持自适应工具路由,在多个数据集上取得优异成绩,证明其为视频理解的有效范式。
AI 中文摘要
从视频中自动理解复杂足球场景对当代视觉语言模型(VLMs)仍是重大挑战,其存在跨模态对齐浅、多步推理和工具集成受限等问题。我们提出TreeSoc,一种结构化推理框架,将足球视频问答重新表述为分层搜索问题而非单步预测。具体而言,它采用动态深度优先搜索机制,将复杂查询分解为有序子任务,支持自适应工具路由。在SoccerBench上,TreeSoc在TextQA、ImageQA和VideoQA上分别达到85.2%、87.4%和82.2%的准确率,在NExT-QA上达到74.16%的准确率,证明了结构化、工具增强的树推理是视频理解的有效范式。
英文摘要
Automated understanding of complex soccer scenarios from video remains a significant challenge for contemporary vision-language models (VLMs), which suffer from shallow cross-modal alignment and exhibit fundamental limitations in multi-step reasoning and coordinated tool integration. We present TreeSoc, a structured reasoning framework that reformulates soccer video question answering as a hierarchical search problem rather than a single-pass prediction. Specifically, TreeSoc employs a dynamic depth-first search (DFS) mechanism that decomposes complex queries into sequentially ordered sub-tasks, enabling iterative reasoning refinement through explicit intermediate states. This tree-structured decomposition naturally supports adaptive tool routing, wherein domain-specific modules are selectively activated and their outputs incorporated at each reasoning node to produce contextually grounded predictions. On SoccerBench, TreeSoc achieves state-of-the-art performance, with accuracies of 85.2%, 87.4%, and 82.2% on TextQA, ImageQA, and VideoQA, respectively. Additionally, TreeSoc further demonstrates strong cross-domain generalization, attaining 74.16% accuracy on NExT-QA. These results establish structured, tool-augmented tree reasoning as an effective paradigm for robust video understanding. Code is available at: https://github.com/thanhnhan29/TreeSoc.
CommentsAccepted to ICMV 2026