arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BenchmarkAnything:智能体驱动的模拟器就绪微架构基准测试构建

BenchmarkAnything: Agent-Driven Construction of Simulator-Ready Microarchitecture Benchmarks

Xiangfeng Sun, Ceyu Xu, Chen Bai, Yuan Xie

arXiv 2610.09447首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; Fudan University(香港科技大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出智能体驱动的工作流,自动将开源仓库转换为模拟器就绪的基准测试,构建大规模套件,以缓解SPEC过拟合,并揭示更真实的架构性能洞察。

AI 中文摘要

基准测试工作负载的选择在计算机体系结构中至关重要,因为它确立了衡量和指导架构创新的标准。然而,几十年来,仅包含数十个工作负载的SPEC基准测试套件一直是学术架构研究的事实标准,并经常被用作主要的评估和优化目标。当如此小的套件被如此重度依赖时,它可能导致架构过拟合;正如我们的研究和先前的研究所表明的,过于狭窄的关注点可能通过高估某些创新而误导设计决策,产生在SPEC基准测试上表现出色但在更广泛、更现实的工作负载上表现不佳的处理器核心。为了缓解这种过拟合,采用一个大型、全面的基准测试套件是自然的解决方案。然而,将软件剥离为模拟器所需的干净、无干扰的二进制可执行文件所需的巨大工程努力,往往使得这一方案极不切实际。在这项工作中,我们证明了AI智能体为这一挑战提供了优雅的解决方案。我们不是手动策划另一个静态基准测试套件,而是引入了一个智能体驱动的工作流程,能够自主地将任意开源代码仓库转换为模拟器就绪的可执行文件。这种自动化方法使工作负载收集具有高度可扩展性,使我们能够快速从公共仓库中收获数百个多样化的应用程序到我们的基准测试套件中。通过将我们智能体生成的套件与SPEC进行对比分析,我们表明它不仅实现了更高保真度的性能评估,还揭示了传统静态套件无法暴露的新颖架构见解。

英文摘要

The selection of benchmark workloads is of paramount importance in computer architecture, as it establishes the yardstick against which architectural innovations are measured and guided. Yet for decades, the SPEC benchmark suites, comprising merely tens of workloads, have been the de facto standard in academic architectural research, where they are frequently treated as a principal evaluation and optimization target. When a suite this small is relied upon so heavily, it risks architectural overfitting; as our research and prior studies demonstrate, an overly narrow focus can mislead design decisions by overvaluing certain innovations, producing cores that excel on SPEC benchmarks yet underperform on broader, realistic workloads. To mitigate this overfitting, adopting a large, comprehensive benchmark suite is the natural solution. However, the immense engineering effort required to strip software into the clean, interference-free binary executables demanded by simulators often makes this highly impractical. In this work, we demonstrate that AI agents provide an elegant solution to this challenge. Rather than manually curating yet another static benchmark suite, we introduce an agent-driven workflow capable of autonomously transforming arbitrary open-source repositories into simulator-ready executables. This automated approach makes workload collection highly scalable, allowing us to rapidly harvest hundreds of diverse applications from public repositories into our benchmark suite. Through a comparative analysis of our agent-generated suite against SPEC, we show that it not only achieves higher-fidelity performance assessments but also uncovers novel architectural insights that traditional, static suites fail to expose.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑