面向多模态认知的分层能量基模型
A Hierarchical Energy-Based Model for Multimodal Cognition
浏览论文内容
中文总结 AI 辅助
该研究提出IM-LEPP多模态能量基模型,扩展单模态LEPP整合视觉与语言,能解释注意现象、复现心理语言学发现,还与Transformer模型形成可证伪对比并提出实验预测。
中文摘要 AI 辅助
我们提出IM-LEPP(集成多模态潜在能量基预测处理),这是一种分层的多模态认知能量基模型,它扩展了先前提出的单模态模型LEPP,以整合视觉与语言。遵循生成式神经网络是认知动力学有效理论的观点,类似于统计力学与热力学的关系,IM-LEPP将认知建模为潜在状态在学习到的能量景观中流动,而非对神经回路的描述。该架构为轮辐式分层结构,基于Lambon Ralph等人的受控语义认知框架,其中针对视觉对象、场景和语言单元的预测编码通路汇聚于以颞前叶为原型的共享非模态中枢。各通路的自身预测受当前中枢状态调节而非被覆盖,在保留通路特定身份的同时,使每个预测反映完整多模态语境。我们表明该架构能对注意现象(如非注意盲视和内克尔立方体双稳态)提供机理解释,其结构还能复现或启发心理语言学中已确立的发现,包括意外性理论、N400/P600事件相关电位(ERP)成分、花园路径重分析,同时在下文预测的轨迹敏感性方面与Transformer语言模型形成可证伪的对比。我们还讨论了相对于大语言模型(LLM)的数据高效语言获取,概述了语义/情景记忆子系统,将该模型置于预测编码、自由能原理、联合嵌入预测架构(JEPA)和分层时间记忆(HTM)的背景中,并提出了验证其核心主张的具体实验预测。
英文摘要
We propose IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical, energy-based model of multimodal cognition that extends a previously proposed single-modality model (LEPP) to integrate vision and language. Following the view that generative neural networks are effective theories of cognitive dynamics, analogous to how statistical mechanics relates to thermodynamics, IM-LEPP models cognition as latent states flowing through learned energy landscapes rather than as an account of neural circuitry. The architecture is a hub-and-spoke hierarchy, grounded in the controlled semantic cognition framework of Lambon Ralph et al., in which predictive-coding pipelines for visual objects, scenes, and linguistic units converge on a shared amodal hub modeled on the anterior temporal lobe. Each pipeline's own prediction is conditioned by, rather than overwritten by, the current hub state, preserving pipeline-specific identity while letting every prediction reflect the full multimodal context. We show this architecture gives a mechanistic account of attentional phenomena such as inattentional blindness and Necker-cube bistability, and that its structure recovers or motivates independently established findings in psycholinguistics, including surprisal theory, the N400/P600 ERP components, and garden-path reanalysis, alongside a falsifiable contrast with transformer language models on trajectory-sensitivity in next-word prediction. We also discuss data-efficient language acquisition relative to LLMs, outline a semantic/episodic memory subsystem, situate the model against predictive coding, the free-energy principle, JEPA, and Hierarchical Temporal Memory, and propose concrete experimental predictions to test its central claims Key Words: predictive processing; predictive coding; energy-based models; diffusion models; effective theory; computational neuroscience.
发表机构
- General Cognitics
机构由 AI 辅助整理,请以论文原文为准。