AI 中文总结
研究边缘平台大语言模型推理问题,提出过渡感知后端调度方法,结合算子特征与后端选择,经实验验证该方法能降低延迟、能量及能量延迟积,减少切换,提升边缘变压器工作负载调度效率。
AI 中文摘要
边缘平台上高效的大语言模型推理不仅受模型大小限制,还受执行后端因形状而异的性能差异影响。静态后端分配无法利用这种变化,而每个算子独立选择会引入高昂的设备和框架切换成本。本文提出一种用于边缘变压器推理的过渡感知后端调度方法,该方法结合当前算子特征与先前选择的后端,在避免不必要过渡的同时保留有益的特定形状选择。从七个变压器模型的全模型推理运行中收集有序跟踪,并在NVIDIA Jetson平台上对PyTorch即时CPU、PyTorch即时CUDA和ONNX Runtime CPU上的四个常见算子类进行基准测试。通过使用从实际后端切换测量的观察到的算子成本和过渡成本进行基于测量的跟踪重放来评估调度策略。支持的算子动态选择,调度范围外的算子保持静态分配。与最佳静态策略相比,过渡感知调度在9584个有序算子实例和278个精确形状组上平均分别将重放延迟、能量和能量延迟积降低了17.4%、14.4%和28.5%。它还减少了相对于算子本地选择的切换。留一模型评估提高了七个保留模型中六个模型的所有三个目标,并提高了所有七个模型的能量。这些结果表明,结合算子形状和后端过渡上下文可以改善边缘变压器工作负载的选择性后端调度。
英文摘要
Efficient large language model (LLM) inference on edge platforms is limited not only by model size, but also by shape-dependent performance differences across execution backends. Static backend assignment cannot exploit this variation, while independent per-operator selection can introduce costly device and framework switches. This paper presents a transition-aware backend dispatch approach for edge transformer inference. The approach combines current operator features with the previously selected backend to preserve beneficial shape-specific choices while avoiding unnecessary transitions. Ordered traces are collected from full-model inference runs of seven transformer models, and four common operator classes are benchmarked across PyTorch eager CPU, PyTorch eager CUDA, and ONNX Runtime CPU on an NVIDIA Jetson platform. The dispatch policies are evaluated through measurement-backed trace replay using observed operator costs and transition costs measured from actual backend switches. Supported operators are selected dynamically, while operators outside the dispatch scope retain a static assignment. Across 9,584 ordered operator instances and 278 exact shape groups, transition-aware dispatch reduces replayed latency, energy, and energy-delay product relative to the best static policy by 17.4%, 14.4%, and 28.5% on average, respectively. It also reduces switching relative to operator-local selection. Leave-one-model-out evaluation improves all three objectives for six of seven held-out models and improves energy for all seven. These results demonstrate that incorporating operator shape and backend-transition context can improve selective backend dispatch for edge transformer workloads.