arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于多个输入同时进行模逆运算的低延迟数据流图

Low latency data-flow graphs for simultaneous modular inversion of many inputs

Tamas Visegrady

arXiv 2607.11337首次发表:更新:

AI 中文总结

研究多个输入同时模逆运算问题,构建最大并行化数据流图,通过特殊调度计算补码乘积,使逆运算后阶段延迟固定,关键路径延迟为⌈log₂N⌉层乘法、一次逆运算和一层最终并行乘法。

AI 中文摘要

蒙哥马利技巧通过分摊单个共享逆运算来加速多个输入的同时模逆运算,但补码乘积的辅助乘法通常按线性、串行形式调度。我们构建了一个最大并行化的数据流图,通过在所有输入乘积及共享逆运算的累加过程中将辅助乘法调度到空闲乘法器时隙来计算所有补码乘积。这种调度确保逆运算后阶段无论N为何值都恰好增加一层乘法延迟,关键路径延迟为⌈log₂N⌉层乘法、一次逆运算和一层最终并行乘法。

英文摘要

Montgomery's trick accelerates simultaneous modular inversion of $N$ inputs by amortizing a single shared inversion, but auxiliary multiplications for complement products are typically scheduled in a linear, serial form. We construct a maximally parallelizable data-flow graph (DFG) that computes all $\overline{x}$ complement~products by scheduling auxiliary multiplications into idle multiplier slots during accumulation of the product of all inputs, and that of the shared inversion. This scheduling ensures the post-inversion phase adds exactly one multiplication layer of latency regardless of $N$, yielding a critical path latency of $\lceil \log_2 N \rceil$ multiply layers, one inversion, and one final parallel multiply layer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑