发表机构
Celabe(赛拉贝)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大脑编码模型预测响应能否用于预测短视频记忆性,通过将视频片段投影到特定皮质空间并用岭回归预测,发现结果取决于数据集,大脑投影在VideoMem数据集表现更佳,且携带特定记忆信号,不是通用先验知识。
AI 中文摘要
大脑编码基础模型能够很好地预测功能磁共振成像对视频、音频和文本的响应,从而赢得了2025年的阿尔戈onauts挑战赛。我们探讨了在不使用扫描仪的情况下获得的预测响应,是否能作为下游人类行为任务(预测短视频的记忆性)的有用特征视角。我们将每个视频片段投影到TRIBE v2预测的皮质空间中,并使用岭回归预测短期记忆性,与匹配的对照进行比较:大脑投影之前模型自身的V-JEPA2视觉主干。结果表明答案取决于数据集。在Memento10k数据集中,主干获胜(斯皮尔曼相关系数为0.594,大脑投影为0.544);在VideoMem数据集中,大脑投影获胜(0.415对0.368,差值为+0.047,95%置信区间为[+0.009,+0.088])。跨数据集转移也继承了这种差异。视频记忆优势并非样本大小的假象,也不仅仅是主干的压缩。预测的大脑特征携带了一个小但真实的记忆信号,主干在一个数据集上错过而在另一个数据集上没有错过:这不是一个通用领域的先验知识,而是一个特定于数据集的表示。一个视觉正交成分定位于腹侧枕颞叶皮质。代码和预测响应数组已发布;源视频和分数未重新分发。
英文摘要
Brain-encoding foundation models predict fMRI responses to video, audio and text well enough to win the Algonauts 2025 challenge. We ask whether their predicted responses, obtained with no scanner, are a useful feature lens for a human-behavior task: forecasting short-video memorability. Each clip is projected into TRIBE v2's predicted cortical space and scored by ridge regression against a matched control, the model's own V-JEPA2 visual backbone taken before the brain projection. The answer is dataset-dependent. Within Memento10k (499 clips) the backbone wins (Spearman 0.594 vs 0.544); within VideoMem (820 clips) the brain projection wins (0.415 vs 0.368). Because the claim is that the ordering reverses, we test the reversal itself: the dataset-by-representation interaction is +0.097, 95% CI [+0.032, +0.160], two-sided bootstrap p=0.001, and over 10 cross-validation seeds the datasets separate completely (0/10 seeds favor the brain projection on Memento10k, 10/10 on VideoMem). Cross-dataset transfer inherits the split: Memento10k->VideoMem the brain projection wins (+0.076); the reverse loses heavily (-0.311). The VideoMem advantage is not a sample-size artifact (it survives matched training size and PCA-then-ridge) and not mere compression (a compressed, heavily regularized or transfer-tuned backbone stays below it). Predicted-brain features thus carry a small but real memorability signal the backbone misses on one dataset and not the other: a dataset-specific representation, not a domain-general prior. A vision-orthogonal component (partial Spearman 0.19, permutation p=2.5e-4) localizes to ventral occipito-temporal cortex, and predicted BOLD dynamics add nothing beyond the time-average because 3-4 samples per clip cannot resolve the sub-second late memorability response. Our pre-specified within-dataset hypothesis returned NO-GO; the reversal is what survived.
Comments11 pages, 3 figures. v4: retitled; adds a dataset-by-representation interaction test (+0.097, 95% CI [+0.032,+0.160], p=0.001) and seed stability for both datasets, replacing the earlier claim that both within-dataset intervals excluded zero (the Memento10k bound was -0.0003); discloses that the pre-specified hypothesis returned NO-GO; adds a regularization-grid check