arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

商树算法:具有有界符号深度和跨子树抵消的延迟除法计算

Quotient Tree Arithmetic: Deferred-Division Computation with Bounded Symbolic Depth and Cross-Subtree Cancellation

Gregory Magarshak

arXiv 2607.22612首次发表:更新:

发表机构

IE University; Safebots Inc.(IE大学; 赛博机器人公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出商树算法(QTA),值以延迟商对表示,比值延迟求值。有界深度增长、跨子树抵消和延迟稳定性三个定理支撑。用于机器学习训练有多种优势,还提出矢量化硬件归一化指令,代数基础是环的局部化,联系代数结构理论与硬件原生算术。

AI 中文摘要

我们引入了商树算法(QTA),这是一种计算框架,其中值表示为延迟商对(N,D),其比值在指定的具体化边界处进行延迟求值。该框架适用于任何领域:将IEEE 754双精度数用作精确整数容器可在2^53精度窗口内实现精确的有理数运算;任意的IEEE双精度数可将覆盖范围扩展到超越值,包括机器学习激活函数如exp(x)和sqrt(x)。QTA有三个结构定理作为支撑。一是有界深度增长:每个算术运算最多使树深度增加1,经过m次运算后树大小为O(m),不会出现组合爆炸。二是跨子树抵消:分子和分母位置出现的子树通过引用标识进行抵消,无需算术运算,包括计算一次并共享的超越值。三是延迟稳定性:在具体化边界处进行一次IEEE除法最多引入二分之一单位的舍入误差,而急切求值会引入O(m)单位的舍入误差。对于机器学习训练,QTA提供了:从结构上防止梯度下溢到零;当中间激活函数引用相同时,通过链式法则磁带折叠实现O(1)成本的梯度计算;共享权重批量压缩,将一批B个示例通过L层的DAG存储从O(BLd)减少到O(L + Bd);当分母已知时,跟踪因子抵消用O(1)试除法代替O(log n)的最大公约数计算。我们提出了一种用于SIMD并行有理对归约的矢量化硬件归一化指令(RatCleanup)。其代数基础是环在其乘法集上的局部化,将QTA与代数结构理论联系起来,同时以硬件原生的IEEE算术为基础。

英文摘要

We introduce Quotient Tree Arithmetic (QTA), a computational substrate in which values are represented as deferred quotient pairs (N, D) whose ratio is evaluated lazily at a designated materialization boundary. The framework applies to any domain: IEEE 754 doubles used as exact integer containers give exact rational arithmetic within the 2^53 exactness window; arbitrary IEEE doubles extend coverage to transcendental values including machine learning activations such as exp(x) and sqrt(x). Three structural theorems underpin QTA. (1) Bounded Depth Growth: each arithmetic operation increases tree depth by at most 1, giving O(m) tree size after m operations with no combinatorial explosion. (2) Cross-Subtree Cancellation: subtrees appearing in both numerator and denominator positions cancel via reference identity without arithmetic, including transcendental values computed once and shared. (3) Deferred Stability: a single IEEE division at the materialization boundary introduces at most one-half ULP of rounding error, versus O(m) ULP for eager evaluation. For machine learning training, QTA provides: structural prevention of gradient underflow to zero; O(1)-cost gradient computation via chain-rule tape collapse when intermediate activations are reference-identical; shared-weight batch compression reducing DAG storage from O(BLd) to O(L+Bd) for a batch of B examples through L layers; and tracked factor cancellation replacing O(log n) GCD with O(1) trial division when denominators are known. We propose a vectorized hardware normalization instruction (RatCleanup) for SIMD-parallel rational pair reduction. The algebraic foundation is the localization of a ring at its multiplicative set, connecting QTA to algebraic structure theory while grounding it in hardware-native IEEE arithmetic.

Comments20 pages, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑