发表机构
Columbia University; University of Pennsylvania(哥伦比亚大学; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于果蝇视觉连接组的可训练架构FlyVision,在多个视觉任务上以远少于ResNet18的参数达到接近或更优的性能,验证了连接组启发的计算可扩展至通用视觉。
AI 中文摘要
生物连接组编码了视觉计算的结构化解决方案,可能为人工视觉提供可复用的归纳偏置。我们围绕FlyVision开发了ConnectomeX,这是一种可训练架构,在跨任务扩展模型容量的同时,保留了并行ON/OFF处理、循环计算和群体级图交互。FlyVision在MNIST上以80,608个参数达到99.34%的准确率,在CIFAR-10上以81,408个参数达到78.03%的准确率。在ImageNet-1K上,FlyVision Base和Large分别以180万和370万参数达到60.79%和66.25%的top-1准确率,而带有学习低频分支的Large local-k7模型达到66.53%,相比之下ResNet18以1170万参数达到69.25%。在22类皮肤病基准上,FlyVision Large以299万参数达到63.78%的准确率和95.28%的宏AUROC。在四类胸部X光摄影中,ImageNet预训练的FlyVision Base和Large分别以133万和297万参数达到92.60%和92.76%的准确率,相比之下ImageNet预训练的ResNet18以1118万参数达到91.56%。BrainAGE通过将共享的ImageNet预训练FlyVision Large编码器应用于每次扫描的24个矢状面、冠状面和轴向切片,并通过置信度调制的高斯投票组合切片级年龄估计,将FlyVision扩展到体积T1加权MRI。在433个保留扫描上,三轴融合实现了5.98年的平均绝对误差和R^2=0.868。在224x224分类任务中,最佳FlyVision配置在ImageNet-1K和皮肤病分类上与ResNet18保持三个百分点以内,并在胸部X光摄影上以显著更少的参数超过它。这些结果表明,保守的连接组信息计算可以从紧凑识别扩展到大规模自然和生物医学视觉。
英文摘要
Biological connectomes encode structured solutions to visual computation that may provide reusable inductive biases for artificial vision. We develop ConnectomeX around FlyVision, a trainable architecture that preserves parallel ON/OFF processing, recurrent computation and population-level graph interaction while scaling model capacity across tasks. FlyVision reached 99.34% accuracy on MNIST with 80,608 parameters and 78.03% on CIFAR-10 with 81,408 parameters. On ImageNet-1K, FlyVision Base and Large reached 60.79% and 66.25% top-1 accuracy with 1.8 and 3.7 million parameters, while a Large local-k7 model with a learned low-frequency branch reached 66.53%, compared with 69.25% for ResNet18 with 11.7 million parameters. On a 22-class skin-disease benchmark, FlyVision Large achieved 63.78% accuracy and 95.28% macro-AUROC with 2.99 million parameters. In four-class chest radiography, ImageNet-pretrained FlyVision Base and Large reached 92.60% and 92.76% accuracy with 1.33 and 2.97 million parameters, compared with 91.56% for ImageNet-pretrained ResNet18 with 11.18 million. BrainAGE extends FlyVision to volumetric T1-weighted MRI by applying a shared ImageNet-pretrained FlyVision Large encoder to 24 sagittal, coronal and axial slices per scan and combining slice-level age estimates by confidence-modulated Gaussian voting. On 433 held-out scans, three-axis fusion achieved a mean absolute error of 5.98 years and R^2 = 0.868. Across the 224x224 classification tasks, the best FlyVision configuration remained within three percentage points of ResNet18 on ImageNet-1K and skin-disease classification and exceeded it on chest radiography with substantially fewer parameters. These results show that a conserved connectome-informed computation can scale from compact recognition to large-scale natural and biomedical vision.
Comments27 pages, 17 figures, 7 tables