arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PerfReasoning:大语言模型在硬件性能上的推理能力如何?

PerfReasoning: How Well Do LLMs Reason on Hardware Performance?

Dan Zhao, Karthikeyan Sankaralingam, Christos Kozyrakis, Qijing Huang

arXiv 2609.04476首次发表:更新:

发表机构

NVIDIA(英伟达公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出PerfReasoning基准,评估大语言模型的硬件性能推理与性能模型代码生成能力,发现闭源模型推理准确率高但代码生成难度大,将发布基准支持可复现评估。

AI 中文摘要

性能建模是硬件设计与软件优化的核心,而构建这类模型需要对计算、数据复用、存储及数据移动进行结构化推理。我们推出PerfReasoning,这一基准测试既将大语言模型(LLM)作为直接性能推理器进行评估,也将其作为分析性性能模型代码的生成器进行评估。给定工作负载、架构及映射规范后,模型会对比不同映射方式并预测片外流量与缓冲区需求。最强的闭源模型在基于推理的问答任务上准确率超过90%,最优的开源权重模型达到82.4%。不过,模型构建难度要大得多:GPT-5.6 Sol的通过率超过80%,但所有其他模型配置的平均通过率低于15%,且不同运行间差异显著。针对特定任务的强化学习(RL)使4B模型的映射推理准确率提升了15.7个百分点,而无反馈的多轮自我修正提示则效果不稳定。PerfReasoning揭示了合理的架构推理与可靠的性能模型构建之间存在差距,我们将公开发布该基准以支持可复现的评估并追踪未来进展。

英文摘要

Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code. Given workload, architecture, and mapping specifications, models compare mappings and predict off-chip traffic and buffer requirements. The strongest closed-source models exceed 90% on reasoning-based Q&A, and the best open-weight model reaches 82.4%. However, model construction is substantially harder: while GPT-5.6 Sol exceeds 80% pass rate, all other model configurations average below 45% and vary markedly across runs. Task-specific RL raises a 4B model's mapping-reasoning accuracy by 15.7 points, whereas feedback-free multi-round self-revision prompting is not reliably effective. PerfReasoning exposes the gap between plausible architectural reasoning and reliable performance-model construction. We will publicly release the benchmark to support reproducible evaluation and track future progress.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑