arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36763cs.NI

SCORAS-MoE:LEO卫星网络中MoE-VLMs的联合压缩与资源自适应部署

SCORAS-MoE: Joint Compression and Resource-Adaptive Deployment of MoE-VLMs in LEO Satellite Networks

  • University of Science and Technology of China(中国科学技术大学)
  • Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

Tong Quan, Yuanlong Wan, Huasen He, Yunpeng Hou, Shuangwu Chen, Xiaofeng Jiang, Jian Yang

AI总结:

SCORAS-MoE提出联合压缩与部署框架,通过扰动感知秩分配和在线调度,在LEO卫星网络中高效部署MoE-VLM,提升精度并加速调度。

AI中文摘要:

在卫星上部署大型视觉语言模型(VLM)能够实现星上数据处理并减少原始数据下行传输。然而,星上推理面临两个资源挑战。有限的星上内存和能量要求进行模型压缩和分布式部署。动态的资源可用性要求在光照、电池电量和通信条件变化时快速做出部署决策。我们提出了SCORAS-MoE,一个用于低地球轨道(LEO)卫星网络中混合专家(MoE)VLM的联合压缩与部署框架。为应对资源有限的问题,SCORAS-MoE测量由低秩近似引起的路由MoE输出的扰动,为更敏感的专家分配更高的秩,并将压缩后的模型分片分布到多颗卫星上进行协同推理。压缩模型生成测量精度和推理能量的配置文件。为适应动态资源,在线调度器在每个时隙选择配置文件组合和分片放置。对于每个候选组合,它将放置问题简化为一个最小成本分配问题,通过匈牙利算法求解,同时枚举组合为当前时隙目标产生最优部署。在Qwen3-VL-30B-A3B-Instruct上的实验表明,基于输出扰动分配秩在激进压缩下特别有效,当专家投影保留其原始参数的30%时,与均匀秩分配相比,平均精度绝对提升3.7%。固定配置文件调度器相比评估的近端策略优化(PPO)和进化基线实现了更高的吞吐量、更少的服务切换和更低的电池影响,分别加速8.7倍和183.5倍。自适应配置文件选择进一步改善了服务质量与能量使用之间的平衡。

英文摘要:

Deploying large vision-language models (VLMs) onboard satellites enables onboard data processing and reduces raw data downlink. However, onboard inference faces two resource challenges. Limited onboard memory and energy require model compression and distributed deployment. Dynamic resource availability requires fast deployment decisions as illumination, battery levels, and communication conditions change. We present SCORAS-MoE, a joint compression and deployment framework for mixture-of-experts (MoE) VLMs in low Earth orbit (LEO) satellite networks. To address limited resources, SCORAS-MoE measures the perturbation of the routed MoE output caused by low-rank approximation, assigns higher ranks to more sensitive experts, and distributes compressed model shards across satellites for cooperative inference. The compressed models yield profiles of measured accuracy and inference energy. To adapt to dynamic resources, the online scheduler selects profile compositions and shard placements in each slot. For each candidate composition, it reduces placement to a minimum-cost assignment problem solved by the Hungarian algorithm, while enumerating the compositions yields the optimal deployment for the current-slot objective. Experiments on Qwen3-VL-30B-A3B-Instruct show that allocating ranks based on output perturbation is particularly effective under aggressive compression, with an absolute gain of $3.7\%$ in mean accuracy over uniform rank allocation when expert projections retain $30\%$ of their original parameters. The fixed-profile scheduler achieves higher throughput with fewer service switches and lower battery impact than the evaluated proximal policy optimization (PPO) and evolutionary baselines, with respective speedups of $8.7\times$ and $183.5\times$. Adaptive profile selection further improves the balance between service quality and energy use.

↑