arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

以视频特征为处理因素的因果推断

Causal Inference with Video Features as Treatments

Kentaro Nakamura, Adam Breuer, Michael H. Crespin, Bryce J. Dietrich, Kosuke Imai

arXiv 2607.06126首次发表:更新:

AI 中文总结

研究以视频特征为处理因素的因果推断问题,核心方法是用深度生成模型及纵向神经网络架构,主要贡献为开发新统计方法,可探究视频视觉特征对观众反应的影响并用于方法基准测试。

AI 中文摘要

我们开发了首个以视频特征为处理因素进行因果推断的统计方法。视频是互联网上最具吸引力的内容形式,核心因果问题是观众反应如何随视频过程中展现的处理因素变化。但标准方法因混淆因素的潜在、高维及动态关联而不适用。我们先用深度生成模型重现视频,利用其内部表征进行因果估计,接着确定动态随机干预下平均潜在结果轨迹可非参数识别,最后提出基于纵向神经网络架构的一致且渐近正态估计器。通过构建含10000个超级马里奥兄弟关卡的新因果推断基准进行实证验证,并应用于2020年美国总统竞选电视广告,发现随时间增加候选人出现概率会提高观众平均评价。该方法能让研究者探究视频中视觉特征及其出现位置对观众反应的影响,并在已知真实因果效应的数据集上对新方法进行基准测试。

英文摘要

We develop the first statistical methodology for causal inference with video features as treatments. Video is the most engaging content modality on the internet. A central causal question is how audience reactions change in response to treatment features that unfold over the course of a video. Unfortunately, standard causal inference methods are not applicable because confounding features are latent, high-dimensional, and dynamically related to both the treatment sequence and the outcome trajectory. To address these challenges, we first reproduce each video using a deep generative model and leverage the model's internal representations as learned, low-dimensional summaries of video content for causal estimation. We then establish that the average potential-outcome trajectory under dynamic stochastic interventions is nonparametrically identified. Lastly, we propose a consistent and asymptotically normal estimator based on a longitudinal neural network architecture. We empirically validate our approach by constructing a new causal inference benchmark consisting of $10{,}000$ Super Mario Bros. levels played by fixed Mario AI agents, where ground-truth causal effects are known by construction. Finally, we apply our method to television advertisements from the 2020 U.S. presidential campaign and find that increasing the probability of a candidate appearing over time leads to higher average viewer evaluations. With the proposed methodology, researchers can ask which visual features, appearing at which points in a video, influence audience responses, while benchmarking new methods against datasets with known ground-truth causal effects.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑