发表机构
College of Artificial Intelligence, Beijing Normal University; Kuaishou Technology(北京师范大学人工智能学院; 快手科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出时序因果驱动评估框架,推导VCD、QCD、PCD指标,经多数据集与模型验证,揭示VLMs生成引导的转变,相关指标可用于多模态生成的源级诊断。
AI 中文摘要
视觉语言模型(VLMs)在复杂图像与视频理解任务中的评估日益增多,但传统指标主要评估最终答案质量,几乎未揭示不同信息源如何影响生成过程。我们提出一种因果时序评估框架,该框架在自回归解码过程中追踪视觉输入、问题文本与生成前缀的演变角色。基于结构因果模型,我们采用干预与后门调整推导了三个按步骤索引的因果驱动指标——视觉因果驱动(VCD)、问题因果驱动(QCD)与前缀因果驱动(PCD),用于表征特定源的生成模式,且无需参考答案。在Qwen3-VL-8B-Instruct上针对MAVIS、LLaVA-Video-178K与MiraData开展的实验,以及在InternVL2-8B上的跨模型验证,均揭示了从早期较强的问题与视觉引导向逐渐依赖生成前缀的一致转变。随机干预验证显示,QCD与PCD相比观测性PMI基线分别降低了34.8%与47.1%的恢复误差。在VLMBias上,前缀-视觉不平衡分数在区分先验驱动与视觉基础生成时,达到0.767的AUROC与0.873的AUPRC。这些结果表明,因果驱动轨迹为多模态生成提供了互补的源级诊断。
英文摘要
Vision-language models (VLMs) are increasingly evaluated on complex image and video understanding tasks, yet conventional metrics primarily assess final-answer quality and reveal little about how different information sources shape the generation process. We propose a causal and temporal evaluation framework that traces the evolving roles of visual input, question text, and generated prefixes during autoregressive decoding. Grounded in a Structural Causal Model, we use interventions and backdoor adjustment to derive three step-indexed causal-drive metrics---Visual Causal Drive (VCD), Question Causal Drive (QCD), and Prefix Causal Drive (PCD)---for characterizing source-specific generation patterns without requiring reference answers. Experiments on Qwen3-VL-8B-Instruct across MAVIS, LLaVA-Video-178K, and MiraData, together with cross-model validation on InternVL2-8B, reveal a consistent transition from stronger early question and visual guidance toward increasing reliance on generated prefixes. Randomized-intervention validation shows that QCD and PCD reduce recovery error over observational PMI baselines by 34.8\% and 47.1\%, respectively. On VLMBias, the prefix--visual imbalance score achieves 0.767 AUROC and 0.873 AUPRC for distinguishing prior-driven from visually grounded generations. These results show that causal-drive trajectories provide complementary source-level diagnostics for multimodal generation.