arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12382cs.LGq-bio.NC

用于从图像序列进行端到端认知地图学习的可微克隆结构因果图

Differentiable Clone-Structured Causal Graphs for End-to-End Cognitive Map Learning from Image Sequences

Arash Nikzad, Sasan Sarbishegi, Ali Dasmeh, Muhammad Asif, Parsa Gharavi, Erik Husom, Sagar Sen, Andrew B. Lehr, Olivier Penacchio, Ana Clemente, Tristan M. Stöber

首次发表
浏览论文内容

中文总结 AI 辅助

研究如何让智能体从图像序列构建世界结构化地图。核心方法是将CSCG重构成可微模块gradCSCG并与VQ-VAE前端耦合,有软发射前向传播和损失平衡机制。主要贡献是证明该方法能在多环境下从视觉输入高精度恢复邻接图,CSCG可作深度学习构建块。

中文摘要 AI 辅助

智能体如何仅从原始感官输入及其自身运动的连续序列构建世界的结构化地图,尤其是在自然变化导致精确感官模式很少重复的情况下?克隆结构因果图算法(CSCG)是一种规范的海马体模型,展示了如何从别名观测中学习可解释的地图。然而,CSCG需要预定义的离散字母表,其期望最大化公式不易与现有神经网络模块结合,阻碍了原始图像序列的端到端处理。我们通过将CSCG重新表述为单个完全可微模块gradCSCG,并将其与学习到的向量量化变分自编码器(VQ-VAE)感知前端耦合来消除这一障碍。软发射前向传播允许地图学习目标反馈到感知中,同时一组损失平衡机制减轻联合训练期间的模块崩溃。我们首先证明,梯度训练通过从高度别名观测中恢复房间拓扑结构,在原始符号网格世界上再现了CSCG的结果。其次,我们表明在MNIST图像序列上地图恢复仍然稳健,在该序列中每次访问一个位置都会产生其分配数字的新采样图像。在四个高度别名的环境中,端到端管道成功地直接从视觉输入中以高边缘精度和召回率揭示了潜在的邻接图。这项工作提供了一个原理证明,即CSCG可以作为深度学习架构中的一个可组合构建块。

英文摘要

How can an agent build a structured map of its world from nothing but an ongoing sequence of raw sensory input and its own movements, especially when natural variation means exact sensory patterns rarely repeat? The Clone-Structured Causal Graph algorithm (CSCG), a normative hippocampus model, shows how an interpretable map can be learned from aliased observations. However, CSCG requires a predefined discrete alphabet, and its expectation-maximization formulation is not easily combined with existing neural network modules, preventing the end-to-end processing of raw image sequences. We remove this barrier by reformulating CSCG as a single, fully differentiable module, gradCSCG, and coupling it to a learned vector-quantized variational autoencoder (VQ-VAE) perceptual front-end. A soft emission forward pass allows the map-learning objective to flow back into perception, while a set of loss-balancing mechanisms mitigates module collapse during joint training. We demonstrate, first, that gradient training reproduces CSCG's results on original symbolic grid worlds by recovering room topology from heavily aliased observations. Second, we show that map recovery remains robust on MNIST image sequences, where each visit to a location yields a newly sampled image of its assigned digit. Across four heavily aliased environments, the end-to-end pipeline successfully uncovers the underlying adjacency graph with high edge precision and recall, directly from visual input. This work provides a proof of principle that CSCG can serve as a composable building block in a deep learning architecture.

发表机构

  • Goethe University Frankfurt(歌德大学法兰克福分校)
  • Max Planck Institute for Human Development(马克斯·普朗克人类发展研究所)
  • Max Planck Institute for Empirical Aesthetics(马克斯·普朗克实验美学研究所)
  • SINTEF(挪威科技工业研究院)
  • University Medical Center Göttingen(哥廷根大学医学中心)
  • Institute of Computer Science and Campus Institute Data Science, University Göttingen(哥廷根大学计算机科学与数据科学研究所)
  • Bridging Research in AI and Neuroscience (brAIN), Computer Vision Center(人工智能与神经科学交叉研究中心(brAIN),计算机视觉中心)
  • Computer Science Department, Universitat Autònoma de Barcelona(巴塞罗那自治大学计算机科学系)
  • Epilepsy Center Frankfurt Rhine-Main, Department of Neurology, Goethe University Frankfurt(歌德大学法兰克福分校莱茵-美因癫痫中心,神经学系)
  • Circulant Labs(循环实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑