arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通用编码计算的学习理论基础:掉队者设置

Learning-Theoretic Foundation for General Coded Computing: The Straggler Setting

Parsa Moradi, Behrooz Tahmasebi, Mohammad Ali Maddah-Ali

arXiv 2608.28910首次发表:更新:

发表机构

University of Minnesota, Twin Cities; Harvard University(明尼苏达大学双城分校; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文从学习理论视角提出通用编码计算(GCC),通过端到端均方误差损失建立其理论基础,在最坏情况和概率两种掉队者机制下证明了GCC的损失收敛速率,拓展了编码计算在机器学习工作负载中的适用性。

AI 中文摘要

编码计算已成为缓解分布式计算系统中掉队工作节点影响的强大范式。然而,现有编码计算方案主要针对多项式求值、矩阵乘法等高度结构化计算的精确恢复而设计,且通常依赖严格的恢复阈值。这些假设极大限制了其在现代机器学习工作负载(尤其是深度神经网络(DNN))中的适用性,DNN的计算通常缺乏刚性代数结构,且在许多应用中仅需精确近似而非精确恢复。为解决这一缺口,本文从学习理论视角重新审视编码计算,引入通用编码计算(General Coded Computing,GCC)。GCC未采用现有代数工具,而是通过自然的端到端均方误差损失公式化编码计算,该损失直接衡量期望计算与其恢复估计之间的差异。通过推导合适的上界,并将编码器和解码器限制在具有温和平滑约束的再生核希尔伯特空间(RKHS)中,我们证明编码器和解码器均可表示为RKHS核函数的线性组合,该表示允许高效计算相应系数。此外,该框架使我们能够在两种互补的掉队者机制下为GCC建立理论性能保证:在最坏情况设置中,当有N个工作节点且最多S个掉队者时,对于标准配置,端到端损失至少以O(S³N⁻³)的速率衰减;我们还研究了概率设置,其中每个工作节点以概率p独立成为掉队者,证明期望损失仍能以O(log₁/p³(N)N⁻³)的速率收敛。

英文摘要

Coded computing has emerged as a powerful paradigm for mitigating the impact of straggling workers in distributed computing systems. However, existing coded-computing schemes are predominantly designed for the exact recovery of highly structured computations, such as polynomial evaluation and matrix multiplication, and typically rely on strict recovery thresholds. These assumptions significantly limit their applicability to modern machine-learning workloads, particularly deep neural networks (DNNs), whose computations generally lack rigid algebraic structure and, in many applications, require only accurate approximations rather than exact recovery. To address this gap, we revisit coded computing from a learning-theoretic perspective and introduce General Coded Computing (GCC). Rather than adopting existing algebraic tools, GCC formulates coded computing through a natural end-to-end mean-squared error loss that directly measures the discrepancy between the desired computations and their recovered estimates. By deriving suitable upper bounds and restricting the encoder and decoder to a reproducing kernel Hilbert space (RKHS) with mild smoothness constraints, we show that both the encoder and decoder admit specific representations as linear combinations of RKHS kernel functions. This representation allows the corresponding coefficients to be computed efficiently. Moreover, this framework enables us to establish theoretical performance guarantees for GCC under two complementary straggler regimes. In the worst-case setting with $N$ worker nodes, and at most $S$ stragglers, we show that the end-to-end loss decays at least at rate $O(S^3N^{-3})$ for standard configurations. We then study a probabilistic setting in which each worker independently straggles with probability $p$. We prove that the expected loss can still converge at rate $O(\log_{1/p}^3(N)N^{-3})$.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑