发表机构
The Hong Kong University of Science and Technology; Huawei Noah’s Ark Lab; The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学; 华为诺亚实验室; 香港科学与技术大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多任务推理模型计算浪费问题,提出软硬件协同设计方法,通过轻量级门控网络预测执行掩码实现任务条件计算跳过,在自动驾驶任务中实验,显著减少计算量、延迟并降低能耗。
AI 中文摘要
多任务推理模型在不同任务间共享单个主干网络,但无论活跃任务是哪个,都会执行相同计算,在无关任务操作上浪费能量和周期。我们发现推理开始前通常可用的任务命令能提供免费信号,可在硬件层面利用其跳过不必要计算。我们提出一种软硬件协同设计方法,其中与主干网络联合训练的轻量级门控网络根据任务输入预测每个tile的二进制执行掩码。每个tile对应固定输出通道组,使掩码tile能零开销跳过。这实现了与任务相关的计算减少,每个命令仅激活所需网络子集,无需改变模型架构或推理管道。我们协同设计了完整系统栈:在稀疏目标下学习硬件对齐tile掩码的命令条件训练过程;指令集架构,其指令携带每个tile的位掩码字段,允许硬件在无软件干预下跳过掩码tile;以及具有可配置并行性、双缓冲内存和原生支持稀疏tile执行的INT8数据路径的tile推理加速器。我们在AMD/Xilinx Alveo U50 FPGA上进行原型设计,并在CARLA自动驾驶模拟器中的闭环视觉运动驾驶任务上进行评估。任务条件稀疏化在保持驾驶质量的同时将FLOP减少66 - 76%。设备上的延迟从9.12毫秒减少51 - 59%至3.74 - 4.44毫秒(加速2.1 - 2.4倍),每次推理的能量从263毫焦降至108 - 128毫焦。
英文摘要
Multi-task inference models share a single backbone across diverse tasks, yet execute identical computation regardless of which task is active - wasting energy and cycles on task-irrelevant operations. We observe that the task command, typically available before inference begins, provides a free signal that can be exploited to skip unnecessary computation at the hardware level. We present a HW/SW co-designed approach in which a lightweight gating network, trained jointly with the backbone, predicts per-tile binary execution masks conditioned on the task input. Each tile corresponds to a fixed group of output channels (the native scheduling granularity of the accelerator), enabling masked tiles to be skipped with zero overhead. This yields a task-dependent reduction in compute, where each command activates only the subset of the network it requires, without changes to the model architecture or inference pipeline. We co-design the full system stack: a command-conditioned training procedure that learns hardware-aligned tile masks under a sparsity objective; an instruction set architecture whose instructions carry per-tile bitmask fields, allowing the hardware to skip masked tiles without software intervention; and a tiled inference accelerator with configurable parallelism, double-buffered memory, and INT8 datapath that natively supports sparse tile execution. We prototype on an AMD/Xilinx Alveo U50 FPGA and evaluate on a closed-loop visuomotor driving task in CARLA autonomous driving simulator. Task-conditional sparsity reduces FLOPs by 66-76% while maintaining driving quality. On-device latency decreases by 51-59%, from 9.12 ms to 3.74-4.44 ms (2.1-2.4x speedup), with energy per inference dropping from 263 to 108-128mJ.
CommentsAccepted at IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)