arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37918cs.CVcs.LG

SYNCR:从模拟中诊断和学习跨视频推理

SYNCR: Diagnosing and Learning Cross-Video Reasoning from Simulation

  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami

AI总结:

SYNCR是基于模拟器的框架,通过共享任务生成器提供跨视频推理的评估与训练数据,诊断多模态大模型缺陷,并通过监督微调显著提升推理能力。

AI中文摘要:

跨视频推理需要对齐事件、匹配身份、比较运动以及整合部分观察。评估这些能力并测试如何改进它们,既需要可靠的标签,也需要有针对性的监督。我们引入了SYNCR,一个基于模拟器的框架,通过共享的任务生成器将这两个需求连接起来。SYNCR构建于Habitat、Kubric和CLEVRER之上,从环境状态中推导答案,并在不重叠的视频上提供4,000个评估问题和15,960个训练问题,涵盖八项跨视频推理任务。视觉消融和人工评估检验了对所提供证据的依赖性和答案的可恢复性。对22个多模态大语言模型的评估揭示了在物理比较和场景整合方面持续存在的困难,而增加模型规模并不能一致地解决这些问题。监督微调将Qwen3-VL-8B在SYNCR上的平均准确率从32.6%提升至61.6%,且提升扩展到这些任务中训练时未出现的任务配置和视频来源。向真实视频的迁移在时间排序方面最为一致:在跨越两个模型家族和两种模型规模的三个检查点上,在构建的Assembly101和Panoptic排序集上准确率提高了9.0-20.5个百分点,并在现有时间推理基准上也有额外提升。这些结果确立了SYNCR作为一个受控环境,用于诊断跨视频推理失败、测试其可学习性,并识别合成监督在何处可迁移。

英文摘要:

Reasoning across videos requires aligning events, matching identities, comparing motion, and integrating partial observations. Evaluating these capabilities and testing how to improve them requires both reliable labels and targeted supervision. We introduce SYNCR, a simulator-grounded framework that connects these two needs through shared task generators. Built on Habitat, Kubric, and CLEVRER, SYNCR derives answers from environment state and provides 4,000 evaluation questions and 15,960 training questions over disjoint videos, spanning eight cross-video reasoning tasks. Visual ablations and human evaluation assess dependence on the supplied evidence and answer recoverability. Evaluation of 22 multimodal large language models reveals persistent difficulties in physical comparison and scene integration that increasing model size does not consistently resolve. Supervised fine-tuning raises Qwen3-VL-8B's average SYNCR accuracy from 32.6% to 61.6%, with gains extending to task configurations and video sources absent from training for those tasks. Transfer to real footage is most consistent for temporal ordering: accuracy improves by 9.0-20.5 percentage points on constructed Assembly101 and Panoptic ordering sets across three checkpoints spanning two model families and two model sizes, with additional gains on existing temporal reasoning benchmarks. These results establish SYNCR as a controlled setting for diagnosing cross-video reasoning failures, testing their learnability, and identifying where synthetic supervision transfers.

↑