arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DART-FL:边缘动态推理需求下的突发感知多任务联邦学习

DART-FL: Burst-Aware Multitask Federated Learning under Dynamic Inference Demand at the Edge

Yiming Xie, Pinrui Yu, Geng Yuan, Xue Lin, Ningfang Mi

arXiv 2608.27713首次发表:更新:

发表机构

Northeastern University; University of Georgia(东北大学; 佐治亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DART-FL是一种感知SLO的多任务联邦学习框架,可动态分配边缘设备的推理与训练资源,在高需求任务突发时段提升其准确率,同时维持长期多任务性能。

AI 中文摘要

边缘智能系统日益要求模型训练与在线推理在资源受限设备上共存,而推理需求会随时间在不同任务间发生大幅变化,这带来两个耦合挑战:必须预留足够计算资源用于推理以维持服务水平目标(SLO),同时剩余训练容量应适配特定任务需求,使高频请求任务在训练过程中更早提升性能。我们提出一种感知SLO、需求驱动的多任务联邦学习框架DART-FL,它可共同适配推理-训练资源分配与任务级训练重点。在每个调度间隔,DART-FL利用推理积压量和已分析服务容量确定推理所需的最小资源分配,随后通过队列感知、受DPP启发的调度器将剩余训练容量分配给各任务,并将所得任务分配映射为动态损失权重,这使得推理需求更高的任务能在更早通信轮次获得更多训练重点。客户端训练共享骨干网络与任务特定头,完整多任务模型通过FedAvg聚合。我们在Stanford Cars和Oxford Flowers 102数据集上,使用合成及源自真实阿里 traces 的工作负载评估DART-FL,结果显示DART-FL可动态适配时变推理需求下的推理-训练资源分配,并将高需求任务的学习进度转向其突发时段,在这些任务被高频请求时提升模型准确率,同时保持相当的长期多任务性能。

英文摘要

Edge intelligence systems increasingly require model training and online inference to coexist on resource-constrained devices, while inference demand can vary substantially across tasks over time. This creates two coupled challenges: sufficient computation must be reserved for inference to maintain service-level objectives (SLOs), while the remaining training capacity should adapt to task-specific demand so that frequently requested tasks can improve earlier during training. We propose an SLO-aware, demand-driven multitask federated learning framework (DART-FL) that jointly adapts the inference-training resource split and task-level training emphasis. At each scheduling interval, DART-FL uses the inference backlog and profiled service capacity to determine the minimum resource allocation required for inference. The remaining training capacity is then distributed across tasks using a queue-aware DPP-inspired scheduler, and the resulting task allocations are mapped to dynamic loss weights. This allows tasks experiencing higher inference demand to receive greater training emphasis in earlier communication rounds. Clients train a shared backbone with task-specific heads, and the complete multitask model is aggregated through FedAvg. We evaluate DART-FL using Stanford Cars and Oxford Flowers 102 under both synthetic and real Alibaba trace-derived workloads. Results show that DART-FL dynamically adapts the inference-training resource split to time-varying inference demand and shifts the learning progress of high-demand tasks toward their burst periods, improving model accuracy when those tasks are frequently requested while maintaining comparable long-term multitask performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑