发表机构
University of Illinois Urbana-Champaign; Indian Institute of Technology Delhi(伊利诺伊大学厄巴纳-香槟分校; 印度理工学院德里分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过模拟4f处理器上的混合傅里叶神经算子,发现随机器件测试低估最坏情况误差,而定向搜索可提供误差下界,强调需用搜索方法评估器件性能。
AI 中文摘要
基于波的处理器件有望为神经算子提供快速、节能的傅里叶层。它们通常在随机采样的器件上进行验证,但使用它们需要知道在制造和对准变化下误差可能变得多大。在一个典型化的数值案例研究中,一个混合傅里叶神经算子将其四个光谱层运行在模拟相干4f处理器上,该处理器具有32个带公差的旋钮,其半宽度具有代表性而非校准值。对于120个模型(四个任务、六种训练方法、五个随机种子),我们将N个随机符合规格器件中的最差者与一个搜索得到的器件进行了比较。在具有一次冻结随机静态误差抽取的确定性模拟器上,搜索器件的留出误差是200个蒙特卡洛器件中最大值的1.08-3.10倍,是1000个器件中最大值的1.06-2.71倍。在20次新的静态抽取下,它仍然在120个模型中的116个中超过了200个随机器件中的最大值。在均匀采样下,每个模型抽取到这样一个器件的概率至多为0.37%(双侧95% Clopper-Pearson),但这并未说明其误差有多大。这种差距在搜索预算下的均匀或Sobol采样、共享旋钮、第二种串扰模型、0.25-2的盒子尺度以及像素级器件模型下持续存在。仅使用随机静态误差训练的模型在搜索器件上的误差达到其标称误差的3.7-39.9倍,而在随机和梯度搜索器件上进行微调在全部20个任务-种子对中给出了六种方法中最低的搜索误差。对于两个换热器量,针对每个量的定向搜索在所有39个模型中超过了1000个随机器件中的最差者,因此也超过了Wilks 95/95极限(59个中的最差者)。对于11个模型的平均压力,没有随机器件超过1%的误差阈值,但搜索器件却超过了。随机测试估计误差超过阈值的频率;最差器件搜索给出了误差可能有多大的下界。
英文摘要
Wave-based processors promise fast, energy-efficient Fourier layers for neural operators. They are usually validated on randomly sampled devices, but using them requires knowing how large their error can become under fabrication and alignment variation. In a stylised numerical case study, a hybrid Fourier neural operator runs its four spectral layers on simulated coherent 4f processors with 32 toleranced knobs, whose half-widths are representative rather than calibrated. For 120 models (four tasks, six training methods, five seeds), we compared the worst of N random in-spec devices with a searched one. On a deterministic simulator with one frozen draw of the random static errors, the searched device's held-out error was 1.08-3.10 times the maximum over 200 Monte Carlo devices and 1.06-2.71 times that over 1000. With 20 fresh static draws, it still exceeded the maximum over 200 random devices in 116 of 120 models. Under uniform sampling, the probability of drawing such a device is at most 0.37% per model (two-sided 95% Clopper-Pearson), which says nothing about how large its error is. The gap persisted with uniform or Sobol' sampling at the search's budget, shared knobs, a second crosstalk model, box scales of 0.25-2 and a pixel-level device model. Models trained only with random static errors reached 3.7-39.9 times their nominal error on searched devices, and fine-tuning on random and gradient-searched devices gave the lowest searched error of the six in all 20 task-seed pairs. For two heat-exchanger quantities, a search targeted at each exceeded the worst of 1000 random devices in all 39 models, and hence the Wilks 95/95 limit (worst of 59). For the mean pressure of 11 models, no random device exceeded a 1% error threshold, but the searched device did. Random testing estimates how often errors exceed a threshold; worst-device search gives a lower bound on how large they can be.
Comments50 pages (19 main text and references, 31 Supplementary Information), 5 figures, 1 table