arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38819cs.CVcs.AIq-bio.NC

未来视频生成比观察视频更符合人类视觉皮层

Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video

Chang-Bae Bang, Hyungjin Chung, Byung-Hoon Kim

首次发表
浏览论文内容

中文总结 AI 辅助

本研究假设未来视频生成的内部表征比观察视频更符合人类视觉皮层的预测性处理,通过fMRI和视频扩散模型验证,并发现人类更偏好此类生成视频。

中文摘要 AI 辅助

研究视觉模型内部表征与视觉皮层对相同观察刺激的反应之间的对齐,使我们能够更好地理解人类视觉处理。然而,迄今为止的研究在很大程度上忽略了一个事实:人类大脑不仅处理观察到的视觉刺激,还会根据已观察到的内容预测即将到来的刺激。据此,我们假设用于生成未来视频帧的内部表征比观察视频本身的表征更符合人类视觉处理的预测性质。为此,我们比较了人类观看视频时的fMRI反应在视觉皮层中的对齐情况,以及来自两种视频扩散模型(一种自回归(AR)模型及其非AR基础模型)的内部表征。我们首先对AR视频扩散模型进行模型内分析,表明用于未来视频生成的表征比观察视频的表征更符合视觉皮层。然后,我们将AR模型的内部表征与其非AR基础模型的内部表征进行比较,再次表明用于未来视频生成的表征比基础模型用于观察视频重建的表征更符合视觉皮层。具体而言,观察视频重建的对齐集中在低级视觉皮层,而未来视频生成的对齐集中在高级视觉皮层。最后,我们在人类行为实验中表明,人类更偏好通过放大与视觉皮层对齐更好的各层贡献而生成的视频。

英文摘要

Studying the alignment between the internal representations of vision models and the responses of the visual cortex to the same observed visual stimuli has enabled us to better understand human visual processing. However, studies so far have largely overlooked the fact that the human brain not only processes observed visual stimuli, but also predicts upcoming stimuli based on what has been observed. Accordingly, we hypothesize that internal representations for generating future video frames are better aligned with the predictive nature of human visual processing than representations of the observed video itself. To this end, we compare the alignment between human video-watching fMRI responses in the visual cortex and the internal representations from two types of video diffusion models, an autoregressive (AR) model and its non-AR base model. We first conduct a within-model analysis of the AR video diffusion model and show that the representations for future video generation align better with the visual cortex than the representations of the observed video. We then compare the internal representations of the AR model with those of its non-AR base model and again show that the representations for future video generation align better with the visual cortex than the representations for observed video reconstruction by the base model. Specifically, the alignment of observed video reconstruction is concentrated in lower-order visual cortex, whereas that of future video generation is concentrated in higher-order visual cortex. Finally, we show in a human behavioral experiment that humans prefer videos generated by amplifying the contributions of individual layers that align better with the visual cortex.

发表机构

  • Yonsei University College of Medicine(延世大学医学院)
  • Institute of Behavioral Sciences in Medicine, Yonsei University College of Medicine(延世大学医学院行为科学研究所)
  • Department of Biomedical Systems Informatics, Yonsei University College of Medicine(延世大学医学院生物医学系统信息学系)
  • Korea University(高丽大学)
  • Yonsei Institute for Digital Healthcare, Yonsei University(延世大学数字医疗研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑