HoliBench:面向CPS-IoT应用中基础模型的跨平台基准测试与部署工具包
HoliBench: A Cross-Platform Benchmarking and Deployment Toolkit for Foundation Models in CPS-IoT Applications
浏览论文内容
中文总结 AI 辅助
HoliBench是一个跨平台基准测试与部署工具包,联合评估基础模型在CPS-IoT设备上的准确性、延迟和能量,通过统一工作流揭示现有工具遗漏的权衡,并支持多模型部署预测。
中文摘要 AI 辅助
基础模型,包括大语言模型、视觉-语言模型和时间序列基础模型,正越来越多地部署在用于CPS和IoT应用的嵌入式及边缘平台上,在这些场景中,能量、延迟和内存与任务准确性同等重要。现有的基准测试工具孤立地评估模型能力,在假设计算资源充足的情况下报告准确性,而硬件性能分析工具则局限于特定平台且相互不兼容。因此,用户缺乏一个统一的工作流程来在异构设备之间做出部署决策。我们提出了HoliBench,一个模块化的基准测试与部署工具包,它能够联合表征从单板计算机到GPU服务器的各平台上的准确性、延迟和能量。其平台抽象层校准了跨设备的测量,该工具包支持多种模型模态、推理引擎、并发级别以及现有的评估框架。一个交互式界面在预先分析一次并在多项研究中复用的设计空间上,提供了考虑约束的配置选择。使用HoliBench,我们对7种设备类型、3种量化级别、8种推理后端和超过30个任务上的20个模型进行了表征,揭示了现有工具遗漏的权衡:量化仅在支持低精度计算的硬件上降低延迟,准确性提升相对于能量消耗呈现收益递减,对于自回归工作负载,平均推理功率在不同输出长度下大致恒定。我们进一步发现,在顺序共驻执行下,单模型性能概况可以组合。在一个多模型CPS部署中,独立概况预测组合管道的延迟和功率的误差分别在1.2%和2.5%以内,从而无需对每个管道配置进行详尽分析即可进行部署探索。我们将HoliBench作为用于基础模型部署感知评估的开源基础设施发布。
英文摘要
Foundation models, including large language models, vision-language models, and time-series foundation models, are increasingly deployed on embedded and edge platforms for CPS and IoT applications, where energy, latency, and memory are as critical as task accuracy. Existing benchmarking tools evaluate model capability in isolation, reporting accuracy assuming sufficient compute, while hardware profiling tools remain platform-specific and mutually incompatible. As a result, users lack a unified workflow for making deployment decisions across heterogeneous devices. We present HoliBench, a modular benchmarking and deployment toolkit that jointly characterizes accuracy, latency, and energy across platforms from single-board computers to GPU servers. Its platform abstraction layer calibrates cross-device measurement, and the toolkit supports multiple model modalities, inference engines, concurrencies, and existing evaluation harnesses. An interactive interface exposes constraint-aware configuration selection over a design space that is profiled once and reused across studies. Using HoliBench, we characterize 20 models across 7 device types, 3 quantization levels, 8 inference backends, and over 30 tasks, surfacing tradeoffs that existing tools miss: quantization reduces latency only on hardware with low-precision support, accuracy gains show diminishing returns relative to energy, and for autoregressive workloads, average inference power is approximately constant across output lengths. We further find that single-model profiles compose under sequential co-resident execution. In a multi-model CPS deployment, standalone profiles predict combined-pipeline latency and power within 1.2% and 2.5%, enabling deployment exploration without exhaustively profiling every pipeline configuration. We release HoliBench as open-source infrastructure for deployment-aware evaluation of foundation models.
发表机构
- University of California, Los Angeles(加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。