arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

低延迟编码张量-矩阵乘法用于分布式信号处理系统

Low-Latency Coded Tensor--Matrix Multiplication for Distributed Signal Processing Systems

Ahmad Tanha, Ali Khalesi

arXiv 2609.30567首次发表:更新:

发表机构

EURECOM; Sorbonne University; IPSA; LINCS Lab(欧莱科; 索邦大学; 巴黎高等航空航天学院; 林克斯实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出张量感知编码计算框架,直接操作子张量实现分布式模式-1张量-矩阵乘法,通过多变量编码降低解码复杂度,显著减少延迟,适用于航空航天等延迟敏感系统。

AI 中文摘要

大规模信号处理系统,包括航空和航天平台,日益依赖张量操作,其分布式执行受到通信、内存和延迟瓶颈的制约。本文提出了一种张量感知的编码计算框架,用于分布式模式-1张量-矩阵乘法,该框架直接对张量子张量进行操作,避免了显式展开并保持了多维结构。对于固定的分区配置,所提出的张量-PolyDot方案在恢复阈值、通信成本、工作端计算和内存需求方面与传统基于矩阵的PolyDot方案相同。然而,通过利用多变量编码,它将融合节点处的高次单变量插值替换为结构化的低维解码。这在不影响任何其他系统指标的情况下,显著降低了解码复杂度和延迟。此外,基于张量的方法改善了内存局部性,并支持跨张量模式的并行解码,使其非常适合现代硬件架构。数值结果显示,解码加速与张量分区因子成正比,在实际设置中可达数量级的提升。这些优势使所提出的框架特别适用于对延迟敏感的航空和航天分布式信号处理系统。

英文摘要

Large-scale signal processing systems, including aeronautical and aerospace platforms, increasingly rely on tensor operations, whose distributed execution is constrained by communication, memory, and latency bottlenecks. In this paper, we propose a tensor-aware coded computation framework for distributed mode-1 tensor--matrix multiplication that operates directly on tensor subtensors, avoiding explicit unfolding and preserving multi-dimensional structure. For a fixed partitioning configuration, the proposed tensor-PolyDot scheme achieves the same recovery threshold, communication cost, worker-side computation, and memory requirements as conventional matrix-based PolyDot schemes. However, by leveraging multivariate encoding, it replaces high-degree univariate interpolation at the fusion node with structured low-dimensional decoding. This leads to a substantial reduction in decoding complexity and latency, without affecting any other system metric. Additionally, the tensor-based approach improves memory locality and enables parallel decoding across tensor modes, making it well-suited to modern hardware architectures. Numerical results show decoding speedups proportional to the tensor partitioning factor, reaching up to an order-of-magnitude improvement in practical settings. These advantages make the proposed framework particularly relevant to latency-critical aeronautical and aerospace distributed signal processing systems.

CommentsTo appear in Proceedings, IEEE International Workshop on Signal Processing Systems (SiPS), November 2nd-4th, 2026, Bordeaux, France

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑