AI 中文总结
研究指出人工智能模型结构单一的问题,追溯其被抛弃的原因,表明Transformer类似海马体非通用皮层。提出异构拓扑网络作为替代,强调在训练前指定模块性,以结构证据为设计输入,为AI架构师提供设计准则。
AI 中文摘要
人工智能研究人员将最先进的模型描述为大规模重复的事物:Transformer,无论用于文本、像素还是语音,其连接方式都相同。神经科学家将皮层描述为马赛克式的——视觉皮层中密集的第4层用于空间编码,运动皮层中较厚的第5/6层用于时间整合——不同的结构解决不同的任务。本文认为这种差距是结构性错误,而非风格问题,并且是可衡量的。一个世纪的细胞结构研究表明,不同的认知功能是由性质不同的结构实现的,而非通过缩放一个模板。卷积神经网络就是该领域自身的证明:局部感受野和层次深度直接编码了这一先验知识,在比后来的架构所需数据少得多的情况下就能实现强大的图像识别。本文追溯了这一教训是如何被抛弃的:“硬件彩票”使Transformer成为阻力最小的路径,而非原则性选择,而常被视为多样性的专家混合模型,实际上是在相同的专家之间划分参数。功能主义分析表明,Transformer最好被理解为海马体形成的功能类似物,而非通用皮层——这与将皮层视为一个巨大的布洛卡区的错误相同,只是现在该领域已经在一个巨型海马体上标准化,并应用于它从未设计用于的任务:听觉、执行门控、工作记忆。本文最后提出了一种替代方案:异构拓扑网络,一种系统的系统,其中不同的模块保持其计算所需的归纳偏差,并通过标准化接口进行通信。这是人工智能架构师的设计准则,而非认知科学:在训练前指定模块性,将结构证据用作设计输入,而非从训练模型的行为中反向工程架构。
英文摘要
AI researchers describe state-of-the-art models as one thing repeated at scale: the Transformer, wired identically for text, pixels, or speech. Neuroscientists describe the cortex as a mosaic - dense Layer 4 in visual cortex for spatial encoding, thick Layers 5/6 in motion cortex for temporal integration - different jobs solved by different structures. This paper argues the gap is a structural error, not a stylistic one, and is measurable. A century of cytoarchitecture, from Brodmann to single-cell Patch-seq, shows distinct cognitive functions are implemented by qualitatively different structures, not by rescaling one template. The convolutional neural network is the field's own proof: local receptive fields and hierarchical depth encoded this prior directly, reaching strong image recognition on far less data than later architectures needed. The paper traces how this lesson was discarded: the "Hardware Lottery" made the Transformer the path of least resistance, not the principled choice, and Mixture-of-Experts, often cited as diversity, in fact partitions parameters among identical experts. A functionalist analysis shows the Transformer is best understood as a functional analog of the hippocampal formation, not a general-purpose cortex - the same mistake as treating cortex as one giant Broca's area, except the field has now standardized on a giant hippocampus, applied to tasks it was never built for: audition, executive gating, working memory. The paper closes with an alternative: a Heterogeneous Topological Network, a System of Systems in which distinct modules keep the inductive bias their computation demands and communicate through standardized interfaces. This is a design discipline for AI architects, not cognitive science: specify modularity before training, using structural evidence as a design input rather than reverse-engineering architecture from a trained model's behavior.
Comments48 pages, 23 figures