Opto-ViT-v2:用于光子近传感器视觉Transformer加速器的抗噪声片上微调
Opto-ViT-v2: Noise-Resilient On-Chip Fine-Tuning for Photonic Near-Sensor Vision Transformer Accelerators
查看机构详情
- Case Western Reserve University(凯斯西储大学)
- New Jersey Institute of Technology(新泽西理工学院)
- Colorado State University(科罗拉多州立大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究针对硅光子加速器支持片上微调的挑战,提出Opto-ViT-v2框架。采用张量低秩分解和梯度累积稀疏分类器,开发噪声模型。实验表明其在光子噪声下能恢复精度,实现高功率效率,达成片上域适应。
中文摘要 AI 辅助
硅光子(SiPh)加速器通过在微环谐振器(MRR)组上进行矩阵乘法,以高吞吐量和能源效率成为视觉Transformer(ViT)推理的有前途平台。将这些平台扩展到支持片上微调仍然具有挑战性,因为反向传播需要大量激活存储、频繁将权重写回MRR以及对器件级噪声的容忍度。我们提出了Opto-ViT-v2,这是第一个用于近传感器SiPh ViT加速器上参数高效微调(PEFT)的框架。我们的张量低秩分解将预训练的光学权重与一小部分可训练的电子因子分开,大大减少了激活存储和权重更新,同时实现了实际的片上训练。我们还引入了梯度累积稀疏分类器,通过一次性top-k梯度掩码冻结低重要性权重,将分类器训练成本降低约40%。我们还开发了第一个用于光子片上训练的系统级噪声模型,在正向和反向传播过程中捕捉MRR串扰、热漂移和激光幅度噪声的影响。使用来自200多个制造的MRR器件进行校准,该模型表明在相同噪声条件下,低秩因子更新比完全微调更稳健。在VTAB-1K(19个任务)和FGVC少样本基准上的实验表明,Opto-ViT-v2在测量的光子噪声下恢复到干净软件精度的0.3%至0.8%以内,同时实现超过100 KFPS/W,为光子边缘视觉系统实现了实际的片上域适应。
英文摘要
Silicon-photonic (SiPh) accelerators have emerged as a promising platform for Vision Transformer (ViT) inference by performing matrix multiplications on microring-resonator (MRR) banks with high throughput and energy efficiency. Extending these platforms to support on-chip fine-tuning remains challenging because backpropagation requires large activation storage, frequent weight write-back to MRRs, and tolerance to device-level noise. We present Opto-ViT-v2, the first framework for parameter-efficient fine-tuning (PEFT) on a near-sensor SiPh ViT accelerator. Our tensorized low-rank decomposition separates pretrained optical weights from a small set of trainable electronic factors (as few as 8K parameters for ViT-Base), greatly reducing activation storage and weight updates while enabling practical on-chip training. We further introduce a gradient-accumulated sparse classifier that freezes low-importance weights through one-shot top-k gradient masking, reducing classifier training cost by about 40 percent. We also develop the first system-level noise model for photonic on-chip training, capturing the effects of MRR crosstalk, thermal drift, and laser amplitude noise during both forward and backward propagation. Calibrated using measurements from more than 200 fabricated MRR devices, the model shows that low-rank factor updates are more robust than full fine-tuning and conventional layer-wise low-rank adaptation under identical noise conditions. Experiments on VTAB-1K (19 tasks) and FGVC few-shot benchmarks demonstrate that Opto-ViT-v2 recovers within 0.3 to 0.8 percent of clean software accuracy under measured photonic noise while achieving more than 100 KFPS/W, enabling practical on-chip domain adaptation for photonic edge vision systems.