ChronoVision:基于潜在状态重构的时序推理
ChronoVision: Temporal Reasoning via Latent State Reconstruction
AI总结:
针对多模态大语言模型时序推理不足的问题,本文提出ChronoVision框架,引入Vbvr-VQA数据集,实验显示其在相关基准上取得最优性能。
AI中文摘要:
多模态大语言模型在被动感知方面表现出色,但在需要多步时序推理的复杂视觉认知任务中表现不佳,这种性能下降主要源于基于语言的推理存在固有歧义,往往无法准确表述连续的视觉变换。为解决该问题,本文提出ChronoVision,这是一款旨在将视觉逻辑与潜在图像对齐的多模态框架。在监督微调阶段,重构视觉头(Reconstructive Visual Head)预测最终变换状态的潜在表示,而ROI注意力定位模块(ROI Attention Locating)通过语义跨度查询引导模型聚焦关键视觉证据;在训练后阶段,本文应用带有隐式过程 grounding 机制的强化学习,该机制由复合奖励函数指导,奖励函数会评估结果正确性、潜在过程对齐以及无监督视觉聚焦。此外,本文引入Vbvr-VQA,这是一款通过将视频推理重新表述为严格的图像排序任务来评估时序跟踪的新型数据集。实验表明,ChronoVision在Vbvr-VQA上达到了最优性能,域内准确率为74.8%、域外准确率为71.6%,同时在极具挑战性的跨域基准IntPhys2上也取得了55.0%的准确率。
英文摘要:
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent imagery. During supervised fine-tuning, a Reconstructive Visual Head predicts the latent representation of the final transformed state, while an ROI Attention Locating module focuses the model on key visual evidence via semantic span queries. In post-training, we apply reinforcement learning with an implicit process grounding mechanism, guided by a composite reward function that evaluates outcome correctness, latent process alignment, and unsupervised visual focus. Furthermore, we introduce Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task. Experiments demonstrate that ChronoVision achieves state-of-the-art performance on Vbvr-VQA with 74.8% in-domain and 71.6% out-of-domain accuracy, alongside a strong 55.0% accuracy on IntPhys2, a highly challenging cross-domain benchmark.