arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于F₂线性代数学习精确的NVIDIA SASS编码器

Learning Exact NVIDIA SASS Encoders with $\mathbb{F}_2$ Linear Algebra

Jiading Gai

arXiv 2608.20532首次发表:更新:

AI 中文总结

针对NVIDIA数据中心GPU缺乏公开SASS汇编器的问题,提出F2Asm系统,基于F₂线性代数学习SASS编码器,支持多款GPU,经测试可精确重汇编SASS指令。

AI 中文摘要

NVIDIA提供了SASS反汇编器,但针对最新数据中心GPU暂无公开的SASS汇编器,限制了可控的机器代码重写。我们提出F2Asm,它从配对的反汇编结果和原始CUBIN指令字中学习精确的128位SASS编码器。据我们所知,F2Asm是首个将SASS指令编码器学习为F₂上向量值仿射映射的系统,也是首个支持Rubin SM107的开源NVIDIA SASS汇编器。F2Asm使用F₂上的高斯消元法逐步构建紧凑基、检测不一致性并拒绝超出学习跨度的输入。它将目标特定的控制位、重定位规则和CUBIN元数据与学习算法分离。我们使用来自官方NVIDIA及第三方生产库、CUDA 13.3包、CUDA 13.4开发者预览存档的3225个CUBIN,训练了针对Hopper SM90/SM90a、Blackwell SM100和Rubin SM107的编码器。往返测试中,F2Asm为每个CUBIN重新汇编反汇编后的SASS,所有对比的可执行文本段均与原始内容完全匹配。

英文摘要

NVIDIA provides a SASS disassembler but no public SASS assembler for recent data-center GPUs, limiting controlled machine-code rewriting. We present F2Asm, which learns exact 128-bit SASS encoders from paired disassembly and original CUBIN instruction words. To our knowledge, F2Asm is the first system to learn SASS instruction encoders as vector-valued affine maps over $\mathbb{F}_2$ and the first open-source NVIDIA SASS assembler to support Rubin SM107. F2Asm uses Gaussian elimination over $\mathbb{F}_2$ to incrementally build a compact basis, detect inconsistencies, and reject inputs outside the learned span. F2Asm separates target-specific control bits, relocation rules, and CUBIN metadata from its learning algorithm. We train encoders for Hopper SM90/SM90a, Blackwell SM100, and Rubin SM107 using 3,225 CUBINs from pinned NVIDIA and third-party production libraries, CUDA 13.3 packages, and CUDA 13.4 Developer Preview archives. In round-trip tests, F2Asm reassembles each CUBIN's disassembled SASS, and all compared executable text sections match the originals exactly. Joint training with F2Asm yields one shared encoder for five Blackwell SM targets and another for three Rubin SM targets, providing strong evidence of a common SASS encoding scheme for instructions shared within each family. Continual training extends the Rubin encoder to all 17,159 previously unsupported cuTile and GROMACS queries with 1,504 additional basis rows, matching the derived lower bound.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑