arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

揭开深度学习编译器前端错误的神秘面纱:一项基于大语言模型辅助的实证研究

Demystifying Deep Learning Compiler Frontend Bugs: An LLM-Aided Empirical Study

Xinyi Yuan, Wei Chen, Jinyi Liu, Pengyu Chen, Jun Wei, Guoquan Wu, Jiaxin Zhu, Tao Huang

arXiv 2607.25651首次发表:更新:

AI 中文总结

对PyTorch 2默认深度学习编译器前端TorchDynamo的fBug展开首次系统实证研究,借助领域知识增强的大语言模型辅助方法,分析fBug并构建分类法,生成测试用例,发现多个新fBug,为深度学习编译器开发和测试提供见解。

AI 中文摘要

深度学习编译器旨在将深度学习程序转换为优化的、特定硬件的代码。通常,其前端将程序转换为基于图的中间表示以实现优化。此阶段引入的缺陷(称为fBug)虽严重但研究不足。我们对PyTorch 2默认的深度学习编译器前端TorchDynamo中的fBug进行首次系统实证研究。利用领域知识增强的大语言模型辅助方法,分析123个fBug并构建分类法。研究结果为开发和测试提供见解,还利用大语言模型生成测试用例,发现了23个先前未知的fBug。

英文摘要

Deep learning compilers (DLCs) are designed to translate deep learning programs into optimized, hardware-specific code. Typically, DLC frontends translate programs into graph-based intermediate representations (IRs) to enable optimizations. Defects introduced during this stage (termed \emph{fBug}s) are severe yet understudied, as prior work predominantly focuses on low-level APIs and operators or treats DLCs as monolithic entities. To bridge this gap, we conduct the first systematic empirical study of \emph{fBug}s in TorchDynamo, the default DLC frontend for PyTorch 2, the most popular DL framework. Leveraging a domain-knowledge-enhanced LLM-aided methodology, we analyze 123 \emph{fBug}s and construct a taxonomy comprising 7 root cause categories and 15 subcategories. Our findings provide actionable insights for DLC development and testing. Furthermore, we leverage the LLM to generate targeted, root cause-aware test cases to detect new bugs. We uncovered 23 previously unknown \emph{fBug}s in recent releases (15 confirmed) across eight (sub)categories, demonstrating the efficacy of our methodology in testing and hardening DLC frontends.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑