arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPEC CPU2026的适配保真度

Adaptation Fidelity of SPEC CPU2026

Doa'a Al-Otoom, Mahesh Madhav

arXiv 2608.27710首次发表:更新:

AI 中文总结

本文首次系统定量分析SPEC CPU2026与其上游开源程序的保真度差距,经单副本延迟与192副本吞吐量运行验证,为SPEC方法学提供数据驱动的验证,明确保真度差距是适配需求的可量化结果。

AI 中文摘要

标准化基准测试常被批评并非“真实工作负载”,但该批评鲜有数据支撑。本文首次对SPEC CPU2026套件与其原始上游开源对应程序之间的“保真度差距”开展系统定量分析。我们编译了SPEC基准测试及其上游应用程序,并在两种场景下使用官方输入工作负载运行:单副本延迟运行和192副本吞吐量运行。研究发现,多数基准测试在单副本运行中表现出高保真度,而少数异常值则揭示了SPEC适配过程的影响。多副本结果进一步凸显了这种适配的必要性:部分基准测试在重负载下的效率显著高于其上游版本,强调了减少I/O的重要性。本研究为SPEC的方法学提供了数据驱动的验证,表明保真度差距并非缺陷,而是为实现可移植性、确定性和以CPU为中心的测量所产生的可量化结果。

英文摘要

Standardized benchmarks are often criticized for not being "real workloads," but this critique is rarely backed by data. This paper provides the first systematic, quantitative analysis of the "fidelity gap" between the SPEC CPU2026 suite and its original, upstream open-source counterparts. We compile both the SPEC benchmarks and their upstream applications and execute them with official input workloads under two scenarios: a single-copy latency run and a 192-copy throughput run. Our findings show that most benchmarks exhibit high fidelity in single-copy runs, while a few outliers reveal the impact of SPEC's adaptation process. The multi-copy results further highlight the necessity of this adaptation: several benchmarks become significantly more efficient than their upstream versions under heavy load, underscoring the importance of I/O reduction. This work offers data-driven validation of SPEC's methodology, showing that the fidelity gap is not a flaw but a quantifiable consequence of enforcing portability, determinism, and CPU-centric measurement.

Comments5 pages, 2 figures, Presented at the 2026 IEEE International Symposium on Workload Characterization (IISWC 2026), Boulder, Colorado

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑