arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10947q-bio.NCcs.LG

柏拉图式大脑桥梁假说:人类大脑网络作为全能模型的结构先验

The Platonic brain bridge hypothesis: human brain networks as an architectural prior for multimodal large language models

发表机构香港科技大学(广州) · 阿里巴巴集团
查看机构详情
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

Pengfei Zhang, Biao Tian, Xiangang Li, Li Liu

首次发表
浏览论文内容

中文总结 AI 辅助

提出柏拉图式大脑桥梁假说,证明全能模型与大脑表征双向对应,并通过Brain-MoE、Brain-AVQA和Brain-Scope三项贡献,将人类大脑网络用作全能模型的结构先验,显著提升性能。

中文摘要 AI 辅助

我们提出柏拉图式大脑桥梁假说:全能模型(omni models)像大脑一样联合处理视频、音频和文本,会收敛于类似大脑的表征,且这种对应关系是双向的。从模型到大脑,七个全能模型的大脑相似性在不同参与者间保持稳定,我们基于其内部隐藏状态的编码模型在Algonauts 2025分布外排行榜上排名第一。从大脑到模型,有三项贡献。Brain-MoE为七个皮层网络各分配一个经大脑预训练的专家,并在全部15个模型-基准组合中将留出准确率平均提高6.42个百分点。Brain-AVQA根据最响应的大脑网络所标记的视频片段构建问题;在域内,真实的网络到专家映射在所有三个模型上均优于打乱后的映射。Brain-Scope使用稀疏自编码器将对应关系定位到一个小子集,移除该子集会削弱所有三种测试基础中的大脑预测。因此,人类大脑网络可作为全能模型的一种可用结构先验。

英文摘要

Multimodal large language models predict brain activity, but brain alignment has been a measurement, not a design tool. We propose the Platonic brain bridge hypothesis: omni models, multimodal large language models that process video, audio and text jointly, converge on brain-like representations usable in both directions. From model to brain, brain-likeness of seven omni models is stable across participants, rises with every input channel in three bases, and our encoders lead the Algonauts 2025 out-of-distribution leaderboard. From brain to model, Brain-MoE fixes the expert partition of a frozen base to the seven networks of human cortex, trains experts on network-labelled Brain-AVQA questions, raises held-out accuracy in all 15 model-benchmark pairs by 6.42 percentage points on average and exceeds capacity-matched random experts in 14. Brain-Scope localizes the correspondence to sparse features whose removal weakens brain prediction. Human brain organization is therefore a usable architectural prior for multimodal large language models.

补充信息

↑