arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

二叉树上的最优(并行)Spooky Pebbling

Optimal (Parallel) Spooky Pebbling on Binary Trees

Mingyu Lee, Sanghyun Lee, Kabgyun Jeong

arXiv 2610.09434首次发表:更新:

发表机构

Seoul National University; Kyung Hee University(首尔国立大学; 庆熙大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对二叉树依赖的双输入计算,提出分块清理策略,给出完全二叉树的最优工作量和并行深度算法,并改进现有界,应用于实际实现可减少约20%的门数量。

AI 中文摘要

Pebble 游戏模拟了在固定空间预算下的计算过程。Spooky pebbling 允许通过测量释放量子内存,并在之后修正由此产生的相位。我们研究了具有二叉树依赖关系的双输入计算,并确定了完全二叉树的最优工作量和并行深度。我们的关键思想是分块清理树,减少中间值的重复重计算。对于具有 $n=2^h-1$ 个顶点且空间预算 $h+1\le s\le n$ 的完全二叉树 $B_h$,我们给出了一种具有渐近最优工作量 $\Theta(nh/\log(s+1))$ 的算法。在最小预算 $s=h+1$ 下,这将 Kornerup、Sadun 和 Soloveichik 的 $O(n\log n)$ 界改进为 $\Theta(n\log n/\log\log n)$,解决了他们的时间最优性问题。我们还构造了一个具有最优深度的并行调度:\\[ \Theta\\!\left(h+\frac ns \max\\!\left\{\frac{h}{\log(s+1)},\\,1+\log^*h\right\}\right). \\] 我们的下界适用于所有完全二叉树。我们还表明,实现最优并行深度可能需要比仅最小化工作量渐近更多的工作量。将我们的调度应用于 Chevignard、Fouque 和 Schrottenloher 公开实现中的 RNS 点加树,在 P-224 实例上减少了 $20.86\\%$ 的 Toffoli/AND 门数量,在 P-256 实例上减少了 $22.66\\%$,同时使用相同的算术电路和峰值工作空间。

英文摘要

Pebble games model computations under a fixed space budget. Spooky pebbling allows quantum memory to be released by measurement, with the resulting phases corrected later. We study two-input computations with binary-tree dependencies and determine the optimal work and parallel depth for complete binary trees. Our key idea is to clean up the tree in blocks, reducing repeated recomputation of intermediate values. For the complete tree $B_h$ with $n=2^h-1$ vertices and every space budget $h+1\le s\le n$, we give an algorithm with asymptotically optimal work $Θ(nh/\log(s+1))$. At the minimum budget $s=h+1$, this improves the $O(n\log n)$ bound of Kornerup, Sadun, and Soloveichik to $Θ(n\log n/\log\log n)$, resolving their time-optimality question. We also construct a parallel schedule with optimal depth \[ Θ\!\left(h+\frac ns \max\!\left\{\frac{h}{\log(s+1)},\,1+\log^*h\right\}\right). \] Our lower bounds hold for every full binary tree. We also show that achieving optimal parallel depth can require asymptotically more work than minimizing work alone. Applying our schedule to the RNS point-addition trees in the public implementation of Chevignard, Fouque, and Schrottenloher reduces their Toffoli/AND gate count by $20.86\%$ for a P-224 instance and $22.66\%$ for a P-256 instance, using the same arithmetic circuits and peak workspace.

Comments21 pages, 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑