发表机构
Mathematical Institute, University of Oxford; Department of Computer Science, University of Oxford; School of Mathematical Sciences, University of Nottingham Ningbo China(牛津大学数学研究所; 牛津大学计算机科学系; 宁波诺丁汉大学数学科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
InfluenceField通过插入干预感知潜在场,实现多模态世界模型中局部干预效应的可识别因果建模,并在CausalVQA上显著提升准确率与鲁棒性。
AI 中文摘要
多模态大语言模型通常能捕捉视觉-语言相关性,但难以预测局部视觉干预如何传播并影响下游答案。我们提出了InfluenceField,一种插入在视觉编码器和语言解码器之间的干预感知潜在场。它将补丁特征提升为连续空间表示,通过共享转移算子传播多步有向影响,并预测局部干预效应。训练过程联合优化了语言建模、跨环境不变性、反事实展开监督和结构正则化。对于非线性有限基总体模型,我们证明目标对齐的干预监督,结合转移上的一步分离条件,将可容许表示限制为位置内重参数化,从而精确恢复完整转移的有向依赖图。线性特化给出了精确的部分覆盖刻画和有限损失稳定性界,场分析推导了系数干预的空间轮廓以及共享通道校准结果。在CausalVQA上,InfluenceField相对于其骨干网络将整体准确率提高了13.1个百分点,在规划和假设类别上提升最大。容量匹配的基线和结构对照将鲁棒性和事实-反事实一致性的提升归因于因果目标而非增加的容量。
英文摘要
Multimodal large language models often capture visual-linguistic correlations but struggle to predict how local visual interventions propagate and affect downstream answers. We introduce InfluenceField, an intervention-aware latent field inserted between the visual encoder and language decoder. It lifts patch features into a continuous spatial representation, propagates directed influence over multiple steps, and predicts local intervention effects through a shared transition operator. Training jointly optimizes language modeling, cross-environment invariance, counterfactual rollout supervision, and structural regularization. For a nonlinear finite-basis population model, we show that target-aligned interventional supervision, together with a one-step separation condition on the transition, restricts admissible representations to within-location reparameterizations, so that the directed dependency graph of the full transition is recovered exactly. A linear specialization gives an exact partial-coverage characterization and a finite-loss stability bound, and the field analysis derives the spatial profile of coefficient interventions together with a shared-channel calibration result. On CausalVQA, InfluenceField improves overall accuracy over its backbone by 13.1 percentage points, with the largest gains on the planning and hypothetical categories. Capacity-matched baselines and structural controls attribute the gains in robustness and factual-counterfactual consistency to the causal objectives rather than to added capacity.
Comments22 pages, 3 figures