arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理前的感知:面向视频理解与问答的动态隐式推理

Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering

Haotian Xia, Zilin Xiao, Junbo Zou, Vicente Ordonez, Hanjie Chen

arXiv 2608.04124首次发表:更新:

发表机构

Rice University; Georgia Institute of Technology(莱斯大学; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视频问答现有方法依赖长文本思维链的问题,提出DyLaR模型,通过动态调整感知与推理隐式变量的使用,在9个视频基准上提升准确率且缩短响应长度。

AI 中文摘要

视频问答要求模型将语言查询与视觉证据建立关联,必要时需跨时间对该证据进行推理。现有方法通常依赖长文本思维链理由,然而许多问题在定位到相关对象、动作或帧后即可回答。我们提出动态隐式推理(Dynamic Latent Reasoning, DyLaR),其首先将问题与一段感知隐式变量(编码查询相关视觉证据的连续隐藏状态)建立关联,随后自适应决定是否在回答前添加推理隐式变量(在隐空间中对该证据进行推理的连续思维)。DyLaR通过将感知隐式变量与已验证视觉证据建立关联、将已验证理由提炼为推理隐式变量,再经强化学习进一步优化推理时机来学习该行为。在9个视频基准和4个多模态语言模型主干上,DyLaR在相同主干的基线模型上提升了平均准确率,且每个查询生成的token少于20个。例如,在Qwen3-VL-4B上,DyLaR将Qwen3-VL-4B-Thinking的平均准确率从54.0提升至58.2,同时将每个查询的响应长度从1220.7个token减少至18.5个token。消融实验进一步表明,关联后的感知隐式变量、受理由监督的推理隐式变量以及自适应路由均能提升准确率。

英文摘要

Video question answering requires models to ground language queries in visual evidence and, when necessary, reason over that evidence across time. Existing methods typically rely on long textual chain-of-thought rationales, even though many questions can be answered as soon as the relevant object, action, or frame is localized. We propose Dynamic Latent Reasoning (DyLaR), which first grounds a question in a short block of perception latents (continuous hidden states that encode query-relevant visual evidence), and then adaptively decides whether to append reasoning latents (continuous thoughts that reason over this evidence in latent space) before answering. DyLaR learns this behavior by grounding perception latents in verified visual evidence and distilling verified rationales into reasoning latents, followed by reinforcement learning that further refines when to reason. Across nine video benchmarks and four multimodal language model backbones, DyLaR improves average accuracy over same-backbone baselines while generating fewer than 20 tokens per query. On Qwen3-VL-4B, for example, DyLaR improves average accuracy over Qwen3-VL-4B-Thinking from 54.0 to 58.2 while reducing response length from 1,220.7 to 18.5 tokens per query. Ablations further show that grounded perception latents, rationale-supervised reasoning latents, and adaptive routing each improve accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑