arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

反应式计算图的成本核算:穷举扫描、顺序变异和反向局部性差距

Cost Accounting for Reactive Computational Graphs: Exhaustive Sweeps, Sequential Mutation, and the Backward-Locality Gap

Abdallah Khemais

arXiv 2607.18323首次发表:更新:

发表机构

ISITCOM, University of Sousse(苏塞大学ISITCOM)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究神经网络计算图穷举干预成本,提出在反应式图引擎上的成本核算方法,包括扫描加速比、变异成本及反向局部性等,通过NeuroDSL验证相关恒等式。

AI 中文摘要

对神经网络计算图进行逐个位点的穷举干预(激活修补扫描、电路发现搜索、系统消融研究)会在每个候选位点变异图,其成本主要由每次变异后的重新计算主导。在一个无效化可证明恰好触及变异节点下游锥的反应式图引擎上,我们给出了此类工作负载的完整成本核算。首先,穷举扫描相对于独立完全重新计算的总加速比不是一个通用常数;其次,我们证明了一系列持久变异的精确成本;第三,我们证明了反向传播中反向传递的正向局部性的精确镜像。每个恒等式都在Julia中的反应式图引擎NeuroDSL上得到验证。

英文摘要

Exhaustive site-by-site interventions on a neural network's computational graph -- activation-patching sweeps, circuit-discovery searches, systematic ablation studies -- mutate the graph at every candidate site, and their cost is dominated by recomputation after each mutation. On a reactive graph engine whose invalidation provably touches exactly the downstream cone of a mutated node, we give a complete cost accounting for such workloads. First, the aggregate speedup of an exhaustive sweep over independent full recomputations is not a universal constant: if per-layer weight varies regularly with depth at Karamata index q, the ratio converges to (q+2)/(q+1) when weight concentrates near the output and to q+2 near the input, recovering 2 only in the depth-uniform case; a wall-clock corollary predicts a ceiling of about 1.79, below 2, until interpreter overhead is compiled away. Second, we prove the exact cost of a sequence of persistent mutations, never undone between insertions: the interleaved cost exceeds the isolated sum by an exact overcount summed over comparable site pairs, with closed-form extremes over insertion orders, while batched application is order-independent and sub-additive, costing exactly the union of the sites' cones plus the fresh nodes. Third, we prove the exact mirror of forward locality for the backward pass, showing it collapses the aggregate speedup to 1 under backpropagation on architectures without long skip connections. Every identity is validated on NeuroDSL, a reactive graph engine in Julia: measured sweep ratios converge to the predicted limits under four cost profiles; the training-mode ratio collapses to 1 at the predicted rate; and all 18 per-graft sequential costs and the batched total match the closed forms at zero tolerance across three insertion orders.

CommentsCompanion paper: "Exact Network Surgery: Functional Invariance and Gradient Plasticity in Reactive Computational Graphs" (submitted concurrently)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑