arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用户辅助协作式分布式推理用于高效的QoS感知自动扩缩容

User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling

Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon, Praveen Kumar Donta

arXiv 2608.11840首次发表:更新:

发表机构

Stockholm University(斯德哥尔摩大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AI推理服务需求增长带来的集中式服务成本上升问题,提出结合专用与用户贡献资源的协作式分布式推理系统,通过高维生成马尔可夫模型优化调度,可在用户规模扩大时提升性能并降低专用资源消耗。

AI 中文摘要

人工智能(AI)推理服务的需求不断增长,这需要可扩展的基础设施,然而集中式服务的成本会随需求上升而增加。我们提出一种协作式分布式推理系统,将专用基础设施与服务用户贡献的资源相结合。专用资源提供维持服务质量(QoS)的基准容量,而自愿贡献的资源可吸收不断增长的需求,无需按比例增加集中式基础设施。为捕捉用户、资源、任务和策略之间的随机动态交互,我们开发了一种具有结构化时间分解的高维生成马尔可夫模型,该模型支持模拟,并为任务调度和QoS感知资源分配优化提供基础。我们针对用户群体、资源容量以及集中式和分布式调度策略对该系统进行评估。模拟结果显示,随着用户群体增长,分布式调度的优势日益凸显,可提升请求完成率和P99延迟,同时大幅降低专用资源消耗。这些结果证明了用户辅助协作式推理在基础设施高效自动扩缩容方面的可行性。

英文摘要

Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. We propose a collaborative distributed inference system combining dedicated infrastructure with resources contributed by service users. Dedicated resources provide baseline capacity for maintaining quality of service (QoS), while volunteered resources absorb increasing demand without proportional growth in centralized infrastructure. To capture stochastic and dynamic interactions among users, resources, tasks, and policies, we develop a high-dimensional generative Markov model with structured temporal factorization. The model supports simulation and provides a foundation for task scheduling and QoS-aware resource allocation optimization. We evaluate the system across user populations, resource capacities, and centralized and distributed scheduling policies. Simulations show that distributed scheduling becomes increasingly advantageous as the user population grows, improving request completion and P99 latency while substantially reducing dedicated resource consumption. These results demonstrate the feasibility of user-assisted collaborative inference for infrastructure-efficient autoscaling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑