arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19387cs.AIcs.AR

AI智能体理解计算机体系结构吗?

Do AI Agents Understand Computer Architecture?

Ambika Sharan, Grigory Chirkov, Soheil Abbasloo

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AutoTuring方法,通过对比命名与匿名参数下的智能体性能,量化AI智能体对计算机体系结构的理解,发现架构知识可提升性能但可被批评循环替代。

中文摘要 AI 辅助

智能体越来越多地被用于设计硬件,并且越来越多地被报道为成功。此类报道证明了设计有所改进;但它们无法证明改进的原因。一个改进加速器的智能体可能是在对机器进行推理,也可能是在对一组其从未理解其含义的参数进行熟练的搜索——而只有前者才能迁移到下一个体系结构。现有的评估无法区分这两种情况,因为它们在保持问题框架不变的同时改变智能体。我们则反其道而行之。AutoTuring将相同的15维加速器空间交给同一个智能体两次:一次以命名的体系结构参数和模拟器计数器呈现,另一次以[0,1]上的匿名变量呈现,同时保持评估器、合法空间和可达最优值相同,因此唯一变化的是问题是否具有意义。两者之间的差距即为测量结果。在包含九个内核的FP16 GEMM基准集上,意义是有回报的:架构师平均比建模的H200高出5.4%,比其盲对应物高出12.3%,且模拟器调用次数减少了70.1%。但这种回报并非独一无二的:一个批评循环为盲智能体恢复了大部分差距,而对架构师则毫无增益,因此体系结构知识和结构化批评表现为替代品而非互补品。我们将这些作为初步发现报告——每个条件下对单个建模加速器进行五到六次运行——并认为比较本身,而非加速器,才是贡献所在。

英文摘要

Agents are increasingly asked to design hardware, and increasingly reported to succeed. Such reports establish that a design improved; they cannot establish why. An agent that improves an accelerator may be reasoning about the machine, or may be searching competently over knobs whose meaning it never recovers -- and only the first transfers to the next architecture. Existing evaluations cannot tell the two apart, because they vary the agent while holding the framing of the problem fixed. We do the opposite. AutoTuring hands the same agent the same 15-dimensional accelerator space twice: once as named architectural knobs with simulator counters, once as anonymous variables on [0,1], with the evaluator, the legal space and the reachable optima held identical, so that the only thing that varies is whether the problem means anything. The gap between the two is the measurement. On a nine-kernel FP16 GEMM basket, meaning pays: the architect beats a modeled H200 by 5.4% and its blind counterpart by 12.3% on average, with 70.1% fewer simulator calls. It does not pay uniquely: a critic loop recovers most of that gap for the blind agent and buys the architect nothing, so architectural knowledge and structured critique behave as substitutes rather than as complements. We report these as preliminary findings -- five to six runs per condition on a single modeled accelerator -- and take the comparison itself, not the accelerator, to be the contribution.

发表机构

  • Microsoft Research(微软研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑