arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HiFi-LLP:用于稳健硬件神经网络架构搜索的具有置信度的高保真、低成本延迟预测器

HiFi-LLP: High-Fidelity, Low-Cost Latency Predictors with Confidence for Robust HW-NAS

Shambhavi Balamuthu Sampath, Behzad Shomali, Nael Fasfous, Moritz Thoma, Judeson Anthony Fernando, Lukas Frickenstein, Pierpaolo Mori, Manoj Rohit Vemparala, Alexander Frickenstein, Walter Stechele

arXiv 2607.11746首次发表:更新:

发表机构

Technical University of Munich; BMW Group; University of Bonn(慕尼黑工业大学; 宝马集团; 波恩大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对硬件感知神经网络架构搜索中硬件在环延迟测量瓶颈及预测不准确问题,提出基于图注意力网络的HiFi-LLP延迟预测器,增加置信度度量,性能优于现有预测器,并构建混合NAS框架,实现加速且保持竞争力。

AI 中文摘要

随着深度神经网络(DNN)越来越多地部署在边缘设备上,硬件感知优化技术,如硬件感知压缩和硬件感知神经网络架构搜索(HW-NAS)变得至关重要。这些方法依赖目标硬件的真实反馈来定制DNN架构以实现高效部署。虽然搜索可并行化,但通过硬件在环(HIL)进行延迟测量因其顺序性仍是瓶颈。近期方法用延迟预测器取代昂贵的HIL反馈,但存在挑战。为此引入HiFi-LLP,一种基于图注意力网络的高保真、低成本延迟预测器,并增加了置信度度量。HiFi-LLP在10%准确率界限上比先前特定平台预测器高出9个百分点,在LatBench数据集中六个设备上Spearman等级相关性高达0.996。还提出混合NAS框架,将低置信度预测路由到HIL,与典型NAS相比实现高达8.6倍加速,同时保持有竞争力的帕累托前沿。

英文摘要

With deep neural networks (DNNs) increasingly deployed on edge devices, hardware (HW)-aware optimization techniques--such as HW-aware compression and HW-aware neural architecture search (HW-NAS)--have become essential. These methods rely on real feedback from the target hardware to tailor DNN architectures for efficient deployment. While the search can be parallelized, latency measurements via hardware-in-the-loop (HIL) remain a bottleneck due to their sequential nature. Recent approaches use latency predictors to replace costly HIL feedback, but challenges persist: (1) platform-specific predictors often require tens of thousands of samples, and (2) inaccurate predictions can mislead the NAS process. To address this, we introduce HiFi-LLP, a high-fidelity, low-cost latency predictor based on graph attention networks, augmented with a confidence metric. HiFi-LLP outperforms prior platform-specific predictors by up to 9 percentage points (p.p.) in the 10% accuracy bound and achieves a Spearman's rank correlation of up to 0.996 across six devices in the LatBench dataset. We further propose a hybrid NAS framework that routes low-confidence predictions to HIL, achieving up to 8.6$\times$ speedup compared to typical NAS while maintaining a competitive Pareto front.

CommentsPublished in the Proceedings of the 2025 IEEE 38th International System-on-Chip Conference (SOCC)

DOI:10.1109/SOCC66126.2025.11235466

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑