arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30413cs.DC

通信感知的模型分布式推理:基于潜在表示压缩

Communication-Aware Model Distributed Inference via Latent Representation Compression

发表机构伊利诺伊大学芝加哥分校 · 东北大学 · 俄亥俄州立大学
查看机构详情
  • University of Illinois Chicago(伊利诺伊大学芝加哥分校)
  • Northeastern University(东北大学)
  • The Ohio State University(俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Peyman Gholami, Theodoros-Thirimachos Davarakis, Teng Li, Miquel Sirera Perelló, Salil Reddy, Ayberk Yarkın Yıldız, Anish Arora, Atilla Eryilmaz, Stratis Ioanni… 展开作者

Peyman Gholami, Theodoros-Thirimachos Davarakis, Teng Li, Miquel Sirera Perelló, Salil Reddy, Ayberk Yarkın Yıldız, Anish Arora, Atilla Eryilmaz, Stratis Ioannidis, Chengzhang Li, Hulya Seferoglu, Ness Shroff

首次发表
浏览论文内容

中文总结 AI 辅助

针对资源受限边缘环境中的分布式模型推理,提出通过潜在表示压缩优化精度与通信成本权衡的框架,并采用闭式解、凸优化及随机对偶下降算法,满足延迟约束并验证有效性。

中文摘要 AI 辅助

我们研究了在资源受限的边缘资源上进行分布式模型推理的优化问题。我们提出了一种框架,通过控制潜在表示压缩来优化模型精度与通信成本之间的权衡,以满足严格的服务质量(QoS)吞吐量目标。对于已知信道状态信息(CSI)的场景,我们推导了单任务情况下的闭式最优解,并将多任务问题转化为一个由每链路注水策略刻画的凸优化问题。我们将这些方法扩展到处理不可预测的环境,采用一种仅依赖因果信道估计的随机对偶下降算法。我们提供了基于Lyapunov的证明,表明我们的方法严格满足长期延迟约束,同时实现有界的最优性差距。我们的结果为在动态、资源受限的分布式系统中最大化流水线式AI任务的性能提供了一个稳健、可扩展的蓝图。我们通过仿真和真实边缘设备的实验验证了我们所提出框架的有效性。

英文摘要

We study optimization of distributed model inference over resource-constrained edge resources. We propose a framework that optimizes the trade-off between model accuracy and communication costs by controlling latent representation compression to meet strict Quality of Service (QoS) throughput targets. For settings with known channel state information (CSI), we derive a closed-form optimal solution for single tasks and reduce the multi-task problem to a convex optimization program characterized by a per-link water-filling strategy. We extend these to handle unpredictable environments via a stochastic dual descent algorithm that relies only on causal channel estimates. We provide Lyapunov-based proofs demonstrating that our approach strictly satisfies long-term delay constraints while achieving a bounded optimality gap. Our results offer a robust, scalable blueprint for maximizing the performance of pipelined AI tasks in dynamic, resource-constrained distributed systems. We verify the effectiveness of our proposed framework through simulations and experiments with real edge devices.

补充信息

↑