arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Iapetus:面向低轨卫星网络中协同ViT推理的内容感知分层调度

Iapetus: Content-Aware Hierarchical Scheduling for Collaborative ViT Inference in LEO Satellite Networks

Yan Chen, Yunxiang Zhang, Guanjun Jiang, Haiquan Wang

arXiv 2609.03318首次发表:更新:

发表机构

Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对低轨卫星网络协同ViT推理的传输开销问题,提出内容感知分层调度器Iapetus,经实验验证其在任务完成率、延迟、能耗上均优于基线方法。

AI 中文摘要

协同推理汇集分布式资源,以在卫星边缘计算中运行计算密集型的视觉Transformer(ViT)。模型分区通过将连续的层组分配给不同节点实现此类协同,但中间激活数据的庞大数据量会产生巨大的传输开销,抵消其带来的益处。令牌压缩可减少下游计算和激活传输,但其质量影响取决于输入内容、模型深度和早期剪枝决策,而层卸载必须适应时变的链路连接和电池状态。我们提出Iapetus(\textbackslash sys),一种内容感知分层调度器,它筛选星座范围内的选项以保留有限的候选集,随后利用质量预测与联合规划将每个候选优化为完整的令牌压缩和层卸载轨迹。统一目标函数平衡单任务延迟、能耗与质量损失,以及累积工作负载和电池压力。我们在NVIDIA Jetson AGX Orin硬件在环测试台上实现Iapetus,并使用其验证后的执行模型,在多个ViT工作负载和星座设置下进行星座规模的轨迹重放。在5任务/秒的速率下,Iapetus完成91.6%的已发布任务,比最强基线MARATD3高出26.1个百分点,同时平均延迟和电池消耗分别降低53.0%和70.8%,且满足质量目标。

英文摘要

Collaborative inference pools distributed resources to run compute-intensive Vision Transformers (ViTs) in satellite edge computing. Model partitioning enables such collaboration by assigning consecutive layer groups to different nodes, but the large volume of intermediate activation data incurs substantial transfer overhead that can erase its benefit. Token compression reduces downstream computation and activation transfer, but its quality impact depends on input content, model depth, and earlier pruning decisions, while layer offloading must adapt to time-varying contact and battery conditions. We present \sys, a content-aware hierarchical scheduler that screens constellation-wide options to retain a bounded candidate set, then refines each candidate into a complete token compression and layer offloading trajectory using quality prediction and joint planning. A unified objective balances per-task latency, energy, and quality loss against accumulated workload and battery pressures. We implement \sys on an NVIDIA Jetson AGX Orin hardware-in-the-loop testbed and use its validated execution model for constellation-scale trace replay across multiple ViT workloads and constellation settings. At \(5\)~tasks/s, \sys accomplishes 91.6\% of released tasks, 26.1 percentage points above MARATD3, the strongest baseline, while reducing mean latency and battery draw by 53.0\% and 70.8\%, respectively, and meeting quality targets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑