Presage:通过智能体引导实验进行预取搜索
Presage: Prefetch Search via Agent-Guided Experiments
浏览论文内容
中文总结 AI 辅助
Presage利用大型语言模型的语义推理,通过智能体引导实验自动插入软件预取,在81个工作负载上实现几何平均10%的性能提升,优于现有方法。
中文摘要 AI 辅助
数据预取是一种成熟的技术,用于缓解缓存未命中延迟并保持处理器持续获得数据。软件处于独特的位置,可以发出预取指令,因为它拥有对当前工作负载的算法知识。然而,插入软件预取是一项耗时、不可预测的任务,并且对微架构或未来代码变更高度敏感。尽管几种基于编译器和基于FDO的方法自动插入和调整软件预取,但它们使用的启发式方法无法泛化,也未考虑大型代码库中出现的问题。在本文中,我们表明对程序的语义理解和推理是插入有效软件预取的关键组成部分。我们还发现,在大型代码库规模上进行预取带来了新的挑战,包括预取相互依赖和干扰。为解决此问题,我们构建了Presage,一个利用大型语言模型的语义推理来插入有效软件预取的系统。Presage通过一个定制的智能体框架处理工作负载,该框架专门设计用于在大型工作负载上导航潜在的预取广阔空间。首先,一个能够访问性能指标的提议智能体识别有前景的软件预取候选。然后,多个长时程优化智能体在循环中优化预取候选,致力于发现性能提升。最后,一个组合智能体处理单独有效预取的组合。通过这种方法,Presage能够插入预取,通过在每个工作负载上探索数十种不同的预取替代方案,在81个工作负载上将性能几何平均提升10%,包括在SPEC CPU 2026上实现2.7%的运行时减少,而先前的SOTA在此失败。与先前SOTA在其评估套件上的比较,Presage实现了几何平均17%的运行时改进,而SOTA为10%。
英文摘要
Data prefetching is an established technique to mitigate cache miss latency and keep the processor saturated with data. Software exists in a unique position to issue prefetches, having algorithmic knowledge of the workload at hand. However, inserting software prefetches is a time consuming, unpredictable task, with a high degree of sensitivity to microarchitecture or future code changes. While several compiler- and FDO-based approaches insert and tune software prefetches automatically, they use heuristics that do not generalize and do not consider issues that arise in large codebases. In this paper, we show that semantic understanding and reasoning about a program is a vital component in inserting effective software prefetches. We also identify that prefetching at the scale of large codebases poses new challenges, including prefetch codependence and interference. To remedy this, we build Presage, a system that leverages the semantic reasoning of Large Language Models to insert effective software prefetches. Presage processes a workload through a custom agent harness specifically designed to navigate the wide space of potential prefetches on large-scale workloads. First, a proposer agent with access to performance metrics identifies promising software prefetching candidates. Then, several long-horizon optimizer agents optimize prefetching candidates in a loop, working towards discovery of performance wins. Finally, a combiner agent handles composition of individually effective prefetches. Through this method, Presage is able to insert prefetches that improve performance by a geomean of 10% across 81 workloads by exploring tens of different prefetching alternatives per workload, including a 2.7% runtime reduction on SPEC CPU 2026 where prior SOTA fails. When compared against prior SOTA on its evaluation suite, Presage achieves a geomean runtime improvement of 17% compared to SOTA's 10%.
发表机构
- Google(谷歌)
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。