AI 中文总结
该研究对比人类与猕猴视觉及相关神经网络,发现预测世界模型兼具跨外观泛化与高神经保真度,提出将运动逐步整合到物体表征是鲁棒动态视觉的关键原理。
AI 中文摘要
智能视觉系统如何在物体外观变化时,将物体的外观特征与运动信息相结合,同时保持鲁棒性?我们通过将人类感知、猕猴下颞叶皮层(IT)的神经活动,与涵盖识别、分割、光流处理和预测世界建模的基于图像和视频的神经网络表征进行比较,来解决这一问题。时间整合提升了物体表征性能,但大多数视频识别模型在外观被干扰、运动结构保留时泛化效果较差,而人类和猕猴IT仍保持鲁棒性。值得注意的是,预测世界模型兼具强大的跨外观泛化能力,且与IT的对应关系最紧密,在神经保真度上优于其他视频建模方法。不过,没有模型能复现从早期外观主导反应到后期外观不变运动编码的皮层转换。这些结果表明,将运动逐步整合到物体表征中是鲁棒动态视觉的一项原理,并指出预测学习是在人工系统中实现该计算的有前景途径。
英文摘要
How does an intelligent visual system combine what objects look like with how they move while remaining robust as appearance changes? We addressed this question by comparing human perception and neural activity in macaque inferior temporal cortex with representations from image- and video-based neural networks spanning recognition, segmentation, optic-flow processing and predictive world modeling. Temporal integration improved object representations, but most video recognition models generalized poorly when appearance was disrupted while motion structure was preserved. Humans and macaque IT remained robust. Notably, predictive world models combined strong cross-appearance generalization with the closest correspondence to IT, outperforming other video-modeling approaches in neural fidelity. Yet no model reproduced the cortical transformation from early appearance-dominated responses toward later appearance-invariant motion coding. These results identify progressive integration of motion into object representations as a principle of robust dynamic vision and implicate predictive learning as a promising route toward realizing this computation in artificial systems.