AI 中文总结
本文针对多租户边缘推理的及时吞吐量优化问题,提出SGPE-SR调度算法,经仿真验证其相比基线方法提升了及时作业完成率。
AI 中文摘要
本文研究面向渐进式边缘推理的因果无线电调度问题,目标是在作业特定截止期限约束下最大化及时推理吞吐量。具体而言,本文构建了一个多租户系统模型,其中每个作业在无线传输与图形处理单元(GPU)计算之间交替执行;该模型捕获了时变上行链路服务、阶段间优先级、感知变体的批处理、两个GPU流上的非抢占式执行以及异构截止期限等特征。为解决无线电-GPU耦合延迟问题,本文提出了分门池证据随机展开(SGPE-SR)调度算法:在一组采样未来路径上选择候选策略与 fallback( fallback 译为“备选”)策略,随后在独立预留集上评估准入决策;在每个预留的候选-备选比较中保留公共随机数,仅当 pooled nominal/recent-history( pooled nominal/recent-history 译为“池化标称/近期历史”)预留优势足够大且独立的近期历史复制统计量非负时,才执行覆盖操作。采用实测GPU性能特征与配对随机实例的系统级仿真表明,在预设的长周期评估中(涉及30个未使用过的种子和3379个提交作业),SGPE-SR算法相比最强的基于速率的基线方法,及时完成率提升了4.27%,每次运行的平均配对增益为3.00个作业,95%自助抽样区间为[1.60, 4.43]。
英文摘要
This paper investigates causal radio scheduling for progressive edge inference, with the goal of maximizing timely inference throughput under job-specific deadlines. Specifically, a multi-tenant system is modeled in which each job alternates between wireless transmission and graphics processing unit (GPU) computation. The model captures time-varying uplink service, inter-stage precedence, variant-aware batching, non-preemptive execution on two GPU streams, and heterogeneous deadlines. To account for delayed radio-GPU coupling, a split-gate pooled-evidence stochastic-rollout (SGPE-SR) scheduler is proposed. Candidate and fallback policies are selected on one set of sampled futures, after which admission is evaluated on an independent held-out set; common random numbers are retained within each held-out candidate-fallback comparison. The override is executed only when its pooled nominal/recent-history held-out advantage is sufficiently positive and an independent recent-history replication statistic is nonnegative. System-level simulations using measured GPU profiles and paired random instances show that, in a prespecified long-horizon evaluation over 30 previously unused seeds and 3,379 offered jobs, SGPE-SR improves timely completions over the strongest rate-based baseline by 4.27%, with a mean paired gain of 3.00 jobs per run and a 95% bootstrap interval of [1.60, 4.43].