arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越缩放:基于经验驱动工作流与经验图记忆的硬件内核优化自演进大语言模型智能体

Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization via an Experience-Driven Workflow and Experience Graph Memory

Siyuan Chen, Runlin Hou, Shenxiu Wu, Yansong Sun, Junming Cao, Yiyu Zhang, Shudi Shao, Junhao Qiu, Zhichao Lu, Qingfu Zhang

arXiv 2608.25570首次发表:更新:

发表机构

City University of Hong Kong; Huawei Technologies Ltd.(香港城市大学; 华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出KOPE框架,通过经验图记忆与主动上下文管理实现硬件内核优化自演进LLM智能体,在相同模型设置下其加速比、通过率等关键指标显著优于现有基线方法。

AI 中文摘要

硬件内核优化需要反复进行编译、正确性测试、性能分析与修订。大语言模型(LLM)智能体可将该过程的部分环节自动化,更强的基础模型、更长的上下文窗口与更长的执行周期已提升了单个任务内的优化效果。但这些进展本身无法让智能体从已完成的优化运行中学习。现有的内核优化智能体很少留存某一决策、其观测到的执行反馈,以及后续使用该证据的决策;保留所有先前轨迹也不切实际,因为不断扩展的历史会与当前任务争夺上下文资源。本文提出KOPE,一个用于硬件内核优化的经验驱动框架。KOPE将带有正确性与性能反馈的优化轨迹记录在经验图记忆中,随后使用主动上下文管理与注入机制,在固定的token预算下检索相关经验。该图保留了决策顺序、观测结果与备选分支,允许在目标硬件上收集的证据为后续优化步骤与任务提供参考。在相同GLM-5.2设置下,KOPE的单算子加速比几何均值是最强竞争基线CANNBot的1.54倍;在包含53个算子的完整消融实验中,主动上下文管理与注入机制使通过率从60.0%提升至84.6%,评估器报告的正域几何均值从0.0382提升至0.0661,优化token消耗从159亿降至11.13亿(相对于被动智能体主导的上下文构建);启用经验图记忆使全套件通过率从55.2%提升至84.6%,有效时序比较的几何均值加速比达1.43倍。这些结果表明,在基础模型保持固定的情况下,通过外部经验支持可实现持续优化。

英文摘要

Hardware kernel optimization requires repeated compilation, correctness testing, profiling, and revision. LLM agents can automate parts of this process, and stronger foundation models, longer context windows, and longer execution horizons have improved optimization within individual tasks. These advances alone do not enable an agent to learn from completed optimization runs. Existing kernel-optimization agents seldom preserve a decision, its observed execution feedback, and the later decisions that use that evidence. Retaining every prior trajectory is also impractical because an expanding history competes with the current task for context. We present KOPE, an experience-driven framework for hardware kernel optimization. KOPE records optimization trajectories with correctness and performance feedback in Experience Graph Memory, then uses Active Context Management and Injection to retrieve relevant experience under a fixed token budget. The graph retains decision order, observed outcomes, and alternative branches, allowing evidence collected on the target hardware to inform later optimization steps and tasks. Under the same GLM-5.2 setting, the geometric mean of KOPE's per-operator speedups is $1.54\times$ that of CANNBot, the strongest competing baseline. In a complete 53-operator ablation, Active Context Management and Injection raises pass rate from 60.0\% to 84.6\%, increases the evaluator-reported positive-field geometric mean from 0.0382 to 0.0661, and reduces optimization token consumption from 15.9B to 1.113B tokens relative to passive agent-led context construction. Enabling Experience Graph Memory raises full-suite pass rate from 55.2\% to 84.6\% and yields a $1.43\times$ geometric-mean speedup on valid timing comparisons. These results support continual optimization through external experience while the foundation model remains fixed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑