AI 中文总结
CNet是一个C++/CUDA复值深度学习框架,利用Wirtinger自动微分和FFT-Hadamard卷积构建物理本位的复值网络,在语言建模、调制分类和相位恢复任务上验证了其有效性。
AI 中文摘要
CNet是一个基于C++/CUDA的框架,用于构建和训练深度复值神经网络(CVNNs),更一般地,用于通过Wirtinger(CR-微积分)导数进行梯度下降来优化复值函数。它采取一种物理本位的立场:网络是对作用于振幅向量的复运算(且通常是酉运算,如DFT)的级联,而分类是一种Born规则测量 $p_k = |z_k|^2 / \\|z\\|^2$,而非对实对数进行softmax。每一层都配有CPU参考实现和经过有限差分校验的CUDA内核,计算图在批次间克隆以用于GPU执行。在基础层之上,我们添加了信号处理原语,将恒等式conv(x,k) = IFFT(FFT(x) · FFT(k))转化为可学习的复卷积网络,同时配备true-Adam优化器和降低内存的推理模式。我们报告了三项研究。首先,一个全复值的、FNet风格的因果序列模型,构建于一种新的$O(N \log N)$因果傅里叶混合器——一种通过Bluestein/chirp-z分解计算的三角掩蔽DFT:经过适当调优后,它在字符级语言建模上达到或超过参数匹配的实值因果FNet,且仅用不到一半的训练步数就达到了实值模型的收敛质量。第二和第三项是瓶颈分析,分别针对无线电调制分类(RML2016.10a)和相干衍射成像的傅里叶相位问题,这些分析精确指出了复值网络仍需要新算子的地方。在所有三项研究中,复值公式都证明能学习到物理上正确的结构。代码:此https URL
英文摘要
CNet is a C++/CUDA framework for building deep complex-valued neural networks (CVNNs) and optimizing complex functions by gradient descent with Wirtinger derivatives. Complex models are underexplored yet natural where data is intrinsically complex -- RF/IQ communications, MRI k-space, radar/SAR, audio spectra -- and phase carries information real networks discard. CNet is physics-native: a network is a cascade of complex (often unitary) operations on an amplitude vector, and classification is a Born-rule measurement p_k = |z_k|^2/||z||^2 rather than a softmax over real logits. Its library of complex layers makes conv(x,k) = IFFT(FFT(x).FFT(k)) learnable via signal-processing primitives -- spectral-padding local kernels (Pad), the inverse DFT, a magnitude nonlinearity |z|^2 (CModulus2), holomorphic powers z^M (CPower), and mean pooling (MeanPool) -- with GPU kernels for every layer, a GPU true-Adam optimizer, and a reduced-memory inference mode. Three studies. (1) A fully complex FNet-style causal language model, built on a new O(N log N) causal Fourier mixer, matches a param-matched real-valued causal FNet on tiny-shakespeare in under half the steps. We introduce Born-rule attention: a content-based causal mixer whose weights are quantum-measurement probabilities |<Q_k,K_j>|^2 of one token against another -- the first attention mechanism built on the Born rule. With rotary position embeddings and per-token complex normalization, a fully complex Born-rule model reaches 1.52 nats/char on tiny-shakespeare, matching softmax attention and beating the Fourier mixer by 0.23. (2,3) Two bottleneck analyses -- radio-modulation classification (RML2016.10a) and the Fourier phase problem of coherent-diffraction imaging -- isolate where CVNNs need new operators. The complex machinery learns the correct structure; open problems are concrete and operator-level. Code: https://github.com/crasmarum/CNet