arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10506cs.ARcs.LGcs.PF

CARB:一种用于CNN推理成本预测与部署筛选的表征引导框架

CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening

  • Florida State University(佛罗里达州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Linh Nguyen, Zhixin Pan

AI总结:

针对资源受限GPU平台上CNN部署前推理成本估计的不足,提出CARB级联混合集成模型与两阶段部署筛选工作流,实现高精度成本预测与高效候选筛选。

AI中文摘要:

随着卷积神经网络(CNN)模型部署在资源受限的GPU平台上,准确的部署前推理成本(能耗、延迟和峰值内存)估计愈发关键。现有方法依赖浮点运算量(FLOPs)、延迟测量或单设备分析作为能耗代理,忽略了架构设计与硬件负载间的非线性交互。我们在RTX 5090和RTX 3080两款GPU平台上,通过GPU遥测技术对13419种CNN配置开展工作负载表征研究,发现能耗、延迟和内存呈现出本质不同的缩放行为:高计算需求下,能耗与延迟的差异达3倍;跨GPU的可迁移性因目标而异——能耗与延迟需平台特定模型,而内存可在两款测试平台间良好迁移。基于这些表征发现,我们开发了CARB,一种级联混合集成模型,可联合预测三个目标,决定系数R²约为0.99;还开发了一种两阶段部署筛选工作流,能在数秒内排除90%以上的候选方案,将庞大的设计空间缩减为经真实硬件验证的帕累托优先短名单。

英文摘要:

Accurate pre-deployment estimation of CNN inference cost--energy, latency, and peak memory--is increasingly critical as models are deployed on resource-constrained GPU platforms. Existing approaches rely on FLOPs, latency measurements, or single-device profiling as energy proxies, overlooking the non-linear interactions between architectural design and hardware load. We present a workload characterization study of 13 419 CNN configurations on two GPU platforms (RTX 5090 and RTX 3080) under GPU telemetry, revealing that energy, latency, and memory exhibit fundamentally distinct scaling behaviors: energy and latency diverge by 3x under high computational demand, and cross-GPU transferability differs by target--energy and latency require platform-specific models while memory transfers well across the two tested platforms. Building on these characterization findings, we develop CARB, a cascade-blended ensemble that jointly predicts all three targets with R2 ~0.99, and a two-stage deployment screening workflow that eliminates over 90% of candidates in seconds, reducing large design spaces to a Pareto-prioritized shortlist validated against real hardware.

↑