预取器为何失效?让智能体来解答
Why Do Prefetchers Fail? Let Agents Answer
浏览论文内容
中文总结 AI 辅助
该研究提出智能体驱动的自动研究流程构建预取器混合模型MoP,在SPEC CPU2006/2017上实现61.1%几何平均IPC加速,超越多款人工设计预取器,为硬件预取器设计提供新路径。
中文摘要 AI 辅助
硬件预取器对处理器性能至关重要,但其设计仍依赖人工投入与专家驱动。架构师检查执行与内存访问轨迹、识别模式、将其转化为在线硬件启发式算法并在仿真中评估,却往往无法保证性能提升。人类专家无法系统检查各类真实工作负载下数十亿指令的轨迹。我们提出一种性能异常驱动的自动研究流程,该流程会反复询问已部署预取器失效的原因,并利用诊断结果构建预取器混合模型(Mixture of Prefetchers, MoP)。每次迭代会将高影响未解释缺失定位到程序计数器,为智能体提供硬件日志、源代码与切片轨迹,通过可运行最小案例验证诊断结果,并为 recurring 模式家族合成专用子预取器。测得的性能与剩余异常会反馈给后续迭代,实现超越模型先验的仿真器在环发现。该研究活动消耗了19.1亿个 DeepSeek V4 Pro 令牌。在 SPEC CPU2006 与 SPEC CPU2017 上,MoP 相较于无预取器实现了61.1%的几何平均 IPC 加速,分别超过人工设计的 Alecto、Berti 与 Pythia 预取器14.5%、21.6%与23.6%。在6nm工艺库中的 RTL 综合显示其片上存储为110 KB,面积为0.0347 mm²。据我们所知,这是首次通过实证证明,智能体驱动的硬件设计流程可生成 RTL 实用的预取器,在未见过的工作负载上超越最先进的人工设计。
英文摘要
Hardware prefetchers are crucial to processor performance, yet their design remains labor-intensive and expert-driven. Architects inspect execution and memory-access traces, identify patterns, translate them into online hardware heuristics, and evaluate them in simulation, often with no guarantee of improvement. Human experts cannot systematically inspect billion-instruction traces across diverse real-world workloads. We present a performance-anomaly-driven autoresearch flow that repeatedly asks why a deployed prefetcher fails and uses the diagnoses to construct the Mixture of Prefetchers (MoP). Each iteration localizes high-impact unexplained misses to program counters, gives agents hardware logs, source code, and sliced traces, validates diagnoses through runnable minimal cases, and synthesizes specialized sub-prefetchers for recurring pattern families. Measured performance and remaining anomalies feed subsequent iterations, enabling simulator-in-the-loop discovery beyond model priors. The campaign consumes 1.91 billion DeepSeek V4 Pro tokens. On SPEC CPU2006 and SPEC CPU2017, MoP achieves a 61.1% geomean IPC speedup over no prefetching, outperforming the human-designed Alecto, Berti, and Pythia prefetchers by 14.5%, 21.6%, and 23.6%, respectively. RTL synthesis in a 6nm library reports 110 KB of on-chip storage and 0.0347 mm^2 area. To our knowledge, this is the first empirical demonstration that an agent-driven hardware-design process can produce an RTL-practical prefetcher that outperforms state-of-the-art human designs on unseen workloads.