arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Rosetta:使用多智能体LLM自动化第一性原理性能建模

Rosetta: Automating First-Principles Performance Modeling Using Multi-Agent LLMs

Karthikeyan Sankaralingam

arXiv 2609.19376首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Rosetta利用多智能体LLM流水线,从论文PDF自动生成第一性原理解析性能模型,通过科学章程、验证-修复循环和集成策略确保质量,在专家评估和作者自评中表现优异。

AI 中文摘要

解析性能模型——从硬件参数推导吞吐量或加速比——使得声明可以独立验证并揭示约束条件,然而由于手工构建一个模型需要数周的专家努力,这类模型很少伴随体系结构论文出现。我们提出了Rosetta,一个多智能体LLM流水线,能够从研究论文PDF自动生成第一性原理解析模型。给定一篇论文作为唯一输入,Rosetta生成数学规范、可执行的Python模型和通俗英语解释——全部自主完成,无需任何人工干预。形式化过程本身是主要价值:它揭示了隐含假设并识别缺失参数。四个设计决策解决了朴素LLM生成中的失败模式:禁止循环推理的科学章程、带有独立评论智能体的验证-修复循环、将功能正确性与科学有效性分离的双重验证,以及利用LLM随机性的最佳N选一集成。我们在三个互补的轨道上评估Rosetta:对12篇里程碑论文的专家评估(CS1)、对97篇未筛选的ISCA 2025和HPCA 2026论文的自动评分(CS2),以及由六个活跃研究小组进行的作者自我评估(CS3)。在CS1中,12篇论文中有10篇的规范质量评分为4-5/5,且没有重大幻觉;在CS2中,56%的通过拟合筛选的论文达到Tier A级洞察质量。最有力的发现来自CS3:Rosetta的输出导致活跃投稿中的声明修订和新实验,六位作者评估者中有五位表示会再次使用它。

英文摘要

Analytical performance models --- derivations of throughput or speedup from hardware parameters --- make claims independently verifiable and expose binding constraints, yet rarely accompany architecture papers because building one by hand takes weeks of expert effort. We present Rosetta, a multi-agent LLM pipeline that automatically generates first-principles analytical models from research paper PDFs. Given a paper as sole input, Rosetta produces a mathematical specification, an executable Python model, and a plain-English interpretation --- all autonomously, with zero human intervention. The formalization process itself is the primary value: it surfaces implicit assumptions and identifies missing parameters. Four design decisions address failure modes of naïve LLM-based generation: a scientific constitution that prohibits circular reasoning, verify-repair loops with independent critic agents, dual verification separating functional correctness from scientific validity, and a best-of-$N$ ensemble that exploits LLM stochasticity. We evaluate Rosetta across three complementary tracks: expert evaluation of 12 landmark papers (CS1), automated scoring of 97 unfiltered ISCA 2025 and HPCA 2026 papers (CS2), and author self-evaluation by six active research groups (CS3). Across CS1, specification quality scores 4--5/5 on 10 of 12 papers with zero significant hallucinations; across CS2, 56\% of fit-screened papers reach Tier~A insight quality. The strongest finding comes from CS3: Rosetta's output led to revised claims and new experiments in active submissions, and five of six author-evaluators said they would use it again.

Comments14 pages, 7 Figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑