神经符号 VideoQA:学习面向真实世界视频问答的组合式时空推理
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
- Harbin University of Science and Technology(哈尔滨理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出神经符号框架 NS-VideoQA,通过 SPN 将视频结构化为符号表示,并用 SRM 与多态执行器完成组合式时空推理,在 AGQA Decomp 上验证了有效性。
AI中文摘要:
组合式时空推理在视频问答(VideoQA)领域构成了一项重大挑战。现有方法难以建立有效的符号推理结构,而这种结构对于回答组合式时空问题至关重要。为应对这一挑战,我们提出了一个名为 Neural-Symbolic VideoQA(NS-VideoQA)的神经符号框架,专门面向真实世界 VideoQA 任务设计。NS-VideoQA 的独特性与优越性体现在两个方面:1)它提出 Scene Parser Network(SPN,场景解析器网络),将静态—动态视频场景转换为 Symbolic Representation(SR,符号表示),对人物、物体、关系和动作年表进行结构化。2)设计了 Symbolic Reasoning Machine(SRM,符号推理机),用于自顶向下的问题分解和自底向上的组合推理。具体而言,框架构建了一个多态程序执行器,以实现从 SR 到最终答案的内部一致推理。因此,我们的 NS-VideoQA 不仅改进了真实世界 VideoQA 任务中的组合式时空推理,还能够通过追踪中间结果进行逐步错误分析。在 AGQA Decomp 基准上的实验评估证明了所提出 NS-VideoQA 框架的有效性。实证研究进一步证实,NS-VideoQA 在回答组合式问题时表现出内部一致性,并显著提升了 VideoQA 任务的时空与逻辑推理能力。
英文摘要:
Compositional spatio-temporal reasoning poses a significant challenge in the field of video question answering (VideoQA). Existing approaches struggle to establish effective symbolic reasoning structures, which are crucial for answering compositional spatio-temporal questions. To address this challenge, we propose a neural-symbolic framework called Neural-Symbolic VideoQA (NS-VideoQA), specifically designed for real-world VideoQA tasks. The uniqueness and superiority of NS-VideoQA are two-fold: 1) It proposes a Scene Parser Network (SPN) to transform static-dynamic video scenes into Symbolic Representation (SR), structuralizing persons, objects, relations, and action chronologies. 2) A Symbolic Reasoning Machine (SRM) is designed for top-down question decompositions and bottom-up compositional reasonings. Specifically, a polymorphic program executor is constructed for internally consistent reasoning from SR to the final answer. As a result, Our NS-VideoQA not only improves the compositional spatio-temporal reasoning in real-world VideoQA task, but also enables step-by-step error analysis by tracing the intermediate results. Experimental evaluations on the AGQA Decomp benchmark demonstrate the effectiveness of the proposed NS-VideoQA framework. Empirical studies further confirm that NS-VideoQA exhibits internal consistency in answering compositional questions and significantly improves the capability of spatio-temporal and logical inference for VideoQA tasks.