发表机构
Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究并发任务的离散计算节点分配问题,提出MARA方法,结合条件流匹配与多智能体策略,在三类工作负载下完成任务率优于基准LARA。
AI 中文摘要
当每个学习任务必须在截止日期前达到目标损失但所需训练量未知时,在并发学习任务间分配有限计算资源是困难的。现有方法将在线损失预测与自适应资源分配相结合,但通常将计算视为可连续分割的吞吐量。我们转而研究一种实际场景:任务随时间到达,计算由离散节点提供,该场景引入了不确定需求和受限的序列决策。我们提出MARA,其通过条件流匹配预测未来损失轨迹,并通过协作多智能体自回归策略协调计算节点;基于势能的进度奖励提供中间训练反馈,同时保留未折现的任务完成目标。在分布内、强化学习和视觉工作负载中,流匹配相较于加权最小二乘降低了剩余资源预测误差;在调度器的训练负载下,MARA平均完成63.46%的任务,比强基准方法LARA(自适应资源分配学习)高出8.54个百分点,且在未见过的更重负载下仍保持领先。
英文摘要
Allocating limited computation among concurrent learning tasks is difficult when each task must reach a target loss before a deadline but its required training effort is unknown. Existing approaches combine online loss prediction with adaptive resource allocation, yet commonly treat computation as continuously divisible throughput. We instead study a practical setting in which tasks arrive over time and computation is provided by discrete nodes. This setting introduces both uncertain demand and constrained sequential decisions. We propose MARA, which predicts future loss trajectories with conditional flow matching and coordinates compute nodes through a cooperative multi-agent autoregressive policy. A potential-based progress reward supplies intermediate training feedback while preserving the undiscounted task-completion objective. Across in-distribution, reinforcement-learning, and vision workloads, flow matching reduces remaining-resource prediction error relative to weighted least squares. At the scheduler's training load, MARA completes 63.46% of tasks on average, 8.54 percentage points above strong baseline Learning with Adaptive Resource Allocation (LARA), and remains ahead under unseen heavier workloads.
Comments10 pages, 4 figures, 6 tables