arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SatDL:联合优化卫星分布式学习中的数据重分布与训练

SatDL: Jointly Optimizing Data Redistribution and Training for Satellite-Based Distributed Learning

Hao Wu, Kin Whye Chew, Yizhan Han, Han Li, Jingxian Wang

arXiv 2608.24516首次发表:更新:

发表机构

School of Computing National University of Singapore Singapore hao\

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对卫星分布式学习中数据非IID导致的训练慢、能耗高问题,提出SatDL框架,通过Distributor-Critic联合优化数据传输与训练,可降低总学习时间和能耗,精度接近现有基准。

AI 中文摘要

基于卫星的分布式学习可利用全球分散的海量传感器数据直接在轨道上训练机器学习模型,避免大规模数据下载至地面服务器。但由于每颗卫星观测不同地理区域,存在严重的非独立同分布(non-IID)数据问题,具体表现为标签不平衡,这会大幅减缓训练收敛速度,延长训练时长,增加太阳能供电卫星的能耗。现有方法要么完全重分布数据以满足独立同分布条件,虽加快收敛但产生大量通信延迟;要么完全不重分布数据,通过修改本地学习算法缓解标签不平衡影响,但仍会延长训练时间并增加能耗。两种极端情况均导致端到端总学习时间(数据传输延迟加训练时间)过长,进而提升星上能耗。本文提出SatDL,一种旨在最小化端到端总学习时间的数据重分布框架,其核心是开发了Distributor-Critic框架,联合建模并优化数据传输延迟与训练时间。通过对1584颗卫星的Starlink星座进行轨迹驱动模拟,以及使用NVIDIA Jetson和A100 GPU在五个数据集上进行硬件仿真评估,结果显示SatDL可将端到端总学习时间最多降低18.6%,星上能耗降低12.23%至88.00%,同时保持推理精度与现有基准仅相差几个百分点。

英文摘要

Satellite-based distributed learning promises to train machine-learning models directly in orbit using massive, globally dispersed sensor data, thereby avoiding large-scale data downloads to ground servers. However, training convergence is significantly slowed by severe non-IID data, specifically label imbalance, as each satellite observes different geographic regions with distinct labels. This imbalance extends training duration and increases energy consumption for solar-powered satellites. Existing approaches either fully redistribute data to enforce IID conditions - accelerating convergence but incurring substantial communication delays - or avoid redistribution entirely by modifying local learning algorithms to mitigate the impact of label imbalance, which, however, still prolong training and increase energy use. Both extremes result in excessive total end-to-end learning time (data-transfer delay plus training time) and thus elevated onboard energy consumption. We present SatDL, a data-redistribution framework designed to minimize total end-to-end learning time. At its core, SatDL develops a Distributor-Critic framework that jointly models and optimizes data-transfer delay and training time. Evaluations through trace-driven simulations of a 1,584-satellite Starlink constellation and hardware emulations using NVIDIA Jetson and A100 GPUs across five datasets show SatDL reduces total end-to-end learning time by up to 18.6% and onboard energy consumption by 12.23-88.00%, while maintaining inference accuracy within a few percentage points of state-of-the-art baselines.

CommentsWithdrawn because this version was submitted prematurely, before all co-authors had completed their review and approved the manuscript for public dissemination. As a result, this version does not represent a manuscript approved by all authors

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑