大语言模型能否对计算机体系结构论文进行深度技术理解?
Can LLMs Perform Technical Comprehension of Computer Architecture Papers?
浏览论文内容
中文总结 AI 辅助
研究大语言模型对计算机体系结构论文的深度理解,通过开源管道Gauntlet进行分析,经与人类分析对比及消融实验,发现其多智能体结构有优势,尤其合成阶段,还将相关资源作为社区资源发布。
中文摘要 AI 辅助
大语言模型能否对计算机体系结构论文进行深度技术理解,不是总结,而是进行结构化批判,指出核心机制、揭示隐含假设并关联超出自身范围的贡献?我们研究了Gauntlet,一个通过五个独立的专家角色评审器和一个对抗性合成阶段来分析论文的开源管道。在20篇ISCA 2025和HPCA 2026论文上,十位研究人员各自撰写分析,然后对他人的论文判断人类分析与Gauntlet分析的优劣。在20次比较中,评估者在15次中更青睐Gauntlet(人类4次,一次平局);在每个分析师的总分上其优势显著(配对Wilcoxon检验,p < 0.01),在关键严谨性方面优势最大,仅在校准方面消失。人类获胜之处在于可信度和实用性而非深度:有自信的错误主张、描述但未讲解的机制或未优先考虑的广度。一项98篇论文的自动消融实验表明,优势来自多智能体结构——该管道在96%的论文上优于作为单一丰富角色智能体运行的相同模型——特别是来自其合成阶段。我们将所有分析、分数和评分标准作为社区资源发布。
英文摘要
Can large language models perform technical comprehension of computer architecture papers--not summarization, but structured critique that names the core mechanism, surfaces buried assumptions, and connects a contribution beyond its own scope? We study Gauntlet, an open-source pipeline that analyzes a paper through five independent expert-persona reviewers and an adversarial synthesis stage. On 20 ISCA 2025 and HPCA 2026 papers, 10 researchers each wrote their own analyses and then judged, for papers other than their own, the human analysis against Gauntlet's. Across the 20 comparisons evaluators preferred Gauntlet in 15 (human in 4, one tie); its advantage is significant on per-analyst totals (two-sided Wilcoxon, p < 0.001) and largest on Critical Rigor. Where humans win, it is on trust and usefulness rather than depth: a confident wrong claim, a mechanism described but not taught, or unprioritized breadth. A 98-paper automated ablation shows the gain comes from the multi-agent structure: the pipeline beats the same model run as a single rich-persona agent on 96% of papers. We release all analyses, scores, and the rubric as a community resource.
发表机构
- University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
- NVIDIA Research(英伟达研究院)
机构由 AI 辅助整理,请以论文原文为准。