用于内存高效的稀疏-二元自组织映射的特征主导码本:在单个消费级GPU上将MEDLINE图谱扩展至105万个神经元
A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPU
浏览论文内容
中文总结 AI 辅助
本文提出特征主导码本的稀疏-二元自组织映射,通过优化码本布局将BMU搜索加速4.5-8.5倍,在单24 GB GPU上训练MEDLINE图谱达105万神经元,是目前最大的自组织映射,性能远超现有方法。
中文摘要 AI 辅助
自组织映射可将大型语料库转换为可浏览的二维图谱,但构建MEDLINE规模的自组织映射一直不切实际:训练中占主导地位的最佳匹配单元(BMU)搜索受限于每轮读取码本所需的带宽。本文表明,该瓶颈在很大程度上是码本布局的人为产物。以特征主导方式存储码本,使每个特征的权重连续(即W[v·M+i]),可将搜索重构为分块稀疏-稠密乘积,其中每个加载的权重列可在一个样本分块中重复使用。仅通过改变布局,在保持实现、精度和更新规则固定的情况下,BMU搜索速度可提升4.5至8.5倍。由于精确argmin的BMU对码本的存储方式不变,该增益无任何代价:在所有图谱尺寸下,保留的量化误差与cuSPARSE基线的偏差均在0.5%以内。相较于该基线,优势是交叉式而非恒定式:在小图谱上速度更快,在128×128时快1.5倍,在256×256时快2.6倍,在512×512时是唯一能在24 GB GPU上运行的方法。结合与半径无关的盒式模糊更新和基于收敛的停止规则,该方法在单个24 GB GPU上,于64×64尺寸时,约72秒内即可在2990万篇MEDLINE文章上训练出收敛的图谱,且可容纳262144个神经元(512×512边),而本文测试的所有替代算法均超出内存限制。在141 GB H200 GPU上,其可达到1048576个神经元(1024×1024边)——据本文所知,这是目前已报道的最大自组织映射。保留误差遵循平滑幂律,在三个数量级的图谱尺寸范围内无拐点,因此分辨率的限制是计算能力而非数据中的任何断点。在相同计算量下,该设计比MedSOM(本文早期MEDLINE图谱背后的CUDA实现)快约82倍,在128×128时比现有最佳多核CPU库快621倍。
英文摘要
Building a self-organising map at MEDLINE scale has been impractical: the best-matching-unit (BMU) search that dominates training is bound by the bandwidth needed to read the codebook every epoch. I show that this bottleneck is largely an artefact of codebook layout. Storing it feature-major with each feature's weights contiguous, W[v.M+i], recasts the search as a tiled sparse-dense product in which every loaded weight column is reused across a tile of samples. Varying only the layout, with implementation, precision and update rule held fixed, accelerates the BMU search by 4.5-8.5x, and because an exact-argmin BMU is invariant to codebook layout, this costs nothing: held-out quantisation error agrees with a cuSPARSE baseline to within 0.5% at every map size. The advantage is a crossover: cuSPARSE$.$SOM is faster at small maps, SparseBin$.$SOM is 1.5x faster at 128x128 and 2.6x at 256x256, and at 512x512 it is the only one that runs on 24 GB without re-engineering its memory path. Paired with a radius-independent box-blur update and a convergence-based stopping rule, it trains a converged map over 29.9 million MEDLINE articles in about 72 s at 64x64 on one 24 GB GPU, and fits 262,144 neurons (512x512) where every alternative I tested exceeds memory; on a 141 GB H200 it reaches 1,048,576 neurons (1024x1024), to my knowledge the largest self-organising map yet reported. Held-out error follows a smooth power law with no elbow across three decades of map size. At matched work, in the configuration benchmarked here, the design is ~82x faster than MedSOM and, at 128x128, 621x faster than the best multicore-CPU library. A post-submission addendum, tuning both implementations symmetrically, accelerates the search a further 5.6-10.1x, brings that 64x64 run to about 13 s, removes the crossover, raises those margins to ~385x and ~3,000x, and narrows two mechanism claims.
发表机构
- College of Medicine and Dentistry, James Cook University(詹姆斯库克大学医学与牙科学院)
机构由 AI 辅助整理,请以论文原文为准。