arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05037cs.ARcs.AI

SparseCraft:面向稀疏计算的智能体软硬件协同优化

SparseCraft: Agentic Hardware-Software Co-Optimization for Sparse Computing

发表机构海得拉巴国际信息技术学院
查看机构详情
  • International Institute of Information Technology, Hyderabad(海得拉巴国际信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

Rajatabha Chakraborty, M P Samartha, Vedant Pahariya, Priyesh Shukla

首次发表
浏览论文内容

中文总结 AI 辅助

SparseCraft利用语言模型在闭环CHIA循环中协同优化稀疏加速器的软硬件设计,通过迭代编辑RTL、内存和调度,在GraphChallenge稀疏DNN层上实现周期、流量和面积的大幅降低。

中文摘要 AI 辅助

稀疏加速器的设计空间通常基于解析模型进行搜索,因此设计点的采纳依据是模型的预测,而非硬件实际行为。SparseCraft通过在封闭的CHIA循环中引入语言模型来弥合这一差距。在15次迭代中,模型读取前一次迭代的实测结果,并通过MCP工具服务器编辑Gemmini加速器的Chisel RTL、内存配置和稀疏内核调度;任何候选方案只有在通过合法性检查、细化、周期精确仿真、与黄金参考逐位比对所有输出以及综合之后才会被计入。该框架将每次测量转化为下一个工作指令,包括诊断出的瓶颈及匹配的策略指导、已尝试设计的历史以及模型自身预测的评分;第二个模型则修复未通过门控的变更。在$512 \ imes 512$的GraphChallenge稀疏DNN层上,该循环相比块稀疏Gemmini基线实现了2.1倍的周期减少、9.8倍的片外流量降低和22.8%的面积缩减,同时建模性能功耗比提升5.61倍,能量延迟积(EDP)降低11.8倍。所利用的杠杆涵盖三个层面:保持稠密操作数驻留的调度消除了9.8倍的流量,模型用Chisel编写的零门控MAC和零行跳过单元降低了能耗,而调整内存大小则缩减了面积。

英文摘要

Sparse-accelerator design spaces are usually searched against analytical models, so a design point is admitted on what a model predicts rather than on what the hardware does. SparseCraft closes that gap with a language model inside a closed CHIA loop. In each of 15 iterations the model reads the measured outcome of the previous one and edits the Chisel RTL, the memory configuration and the sparse-kernel schedule of a Gemmini accelerator through MCP tool servers, and no candidate counts until it has been checked for legality, elaborated, simulated cycle-accurately, checked bit-for-bit on every output against a golden reference, and synthesised. The harness turns each measurement into the next work order, a diagnosed bottleneck with matching strategy guidance, the history of tried designs and a score of the model's own prediction, and a second model repairs changes that fail a gate. On a $512 \times 512$ GraphChallenge sparse-DNN layer the loop reaches 2.1x fewer cycles, 9.8x less off-chip traffic and 22.8% less area than the block-sparse Gemmini baseline, with 5.61x higher modelled perf/W and 11.8x lower EDP. The levers span three layers: a schedule that keeps the dense operand resident removes 9.8x of the traffic, a zero-gated MAC and a zero-row skip unit that the model wrote in Chisel cut energy, and resizing the memories cuts area.

补充信息

↑