发表机构
East China Normal University; Hefei National Laboratory; Shanghai Research Center for Quantum Sciences; Shanxi University(华东师范大学; 合肥国家实验室; 上海量子科学研究中心; 山西大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种薄膜铌酸锂上的有符号非相干光子矩阵核,支持原位反向传播,实现8位乘法与10位累加精度,并成功执行BERT-mini负载,为大规模高效光子神经形态计算提供可扩展基础模块。
AI 中文摘要
人工智能工作负载日益需要结合高吞吐量、高能效和物理可扩展性的计算架构。本文在薄膜铌酸锂上提出了一种原生支持有符号输入和权重的有符号非相干光学矩阵乘法器。我们在16×16光子核上演示了闭环原位反向传播,包括前向计算、非线性操作、误差传播和梯度计算。实测物理输出直接参与优化,从而将实际器件响应和非理想性纳入训练。该核在256个通道上实现了8位乘法和10位累加精度,在拼接的256×256矩阵计算中保持10位精度,并进一步在BERT-mini工作负载中执行Transformer线性操作。这项工作提供了一个基础性、可扩展的构建模块,解决了实际光学计算中的关键限制,从而为大规模、高效率的光子神经形态系统开辟了途径。
英文摘要
Artificial intelligence workloads increasingly demand computing architectures combining high throughput, energy efficiency, and physical scalability. Here we present a signed incoherent optical matrix multiplier on thin-film lithium niobate that natively supports signed inputs and weights. We demonstrate closed-loop in situ backpropagation on a 16 X 16 photonic core, including forward computation, nonlinear operations, error propagation, and gradient computation. Measured physical outputs directly participate in optimization, thereby incorporating actual device responses and nonidealities into training. The core achieves 8-bit multiplication and 10-bit accumulation precision across 256 channels, maintains 10-bit accuracy in tiled 256 X 256 matrix computation, and further executes Transformer linear operations in a BERT-mini workload. This work presents a fundamental, scalable building block that addresses key limitations in practical optical computing, thereby opening avenues for large-scale, high-efficiency photonic neuromorphic systems.