arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18254cs.AIcs.LGcs.PL

无需重新训练的跨方言泛化:MLIR 的模式派生约束解码基准测试与评估

Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR

Plawan Kumar Rath

首次发表
浏览论文内容

中文总结 AI 辅助

研究在 MLIR 中能否用从方言操作定义规范派生的推理时先验替代梯度适应,发布跨三种方言的基准测试集,构建三层模式派生约束栈,实验表明在部分方言上新模式能让模型在速度和验证有效率上超越其他模型,并发布相关资源。

中文摘要 AI 辅助

多级中间表示(MLIR)是现代机器学习编译器基础设施(TensorFlow、JAX/StableHLO、PyTorch Inductor、IREE)的基础,但在代码语言模型预训练语料库中出现的数量很少。MLIR 设计上具有可扩展性,每个应用领域都会推出新方言,因此为每个方言微调模型不可行。研究人员探讨能否从每个方言的操作定义规范(ODS)机械派生的推理时先验替代基于梯度的适应。首先发布了跨三种方言的四个自然语言到 MLIR 的基准测试集,包括 MLIR-Spec-150、Linalg-Spec-30、StableHLO-Spec-30 和 StableHLO-Held-Out-200,共 410 个范围内的自然语言到 MLIR 对,以及一个 25 个程序的语法外压力集和一个人工编写的 n=30 的功能参考集,并在 Apache-2.0 许可下发布,附带 Gebru 数据表和 Croissant 1.0 元数据。其次构建了一个三层模式派生约束栈:基于操作签名的控制流图(C1)、从 ODS 提取的类型格派生的类型域分割(C2)以及驱动五次重试拒绝采样的 SSA 范围验证器(C3)。从 arith+func+memref+linalg 移植到 StableHLO 无需新的约束层代码。在验证器语义由结构约束主导的方言上,模式派生先验使 SmolLM2-1.7B 在每代速度提高 8 到 25 倍的情况下,匹配或超过 15B - 34B 的代码语言模型:在 linalg 方言上,SmolLM2 的验证有效率达到 80.0%(三个种子的平均值,n = 125),比 CodeLlama-34B、Granite-Code-34B 和 StarCoder2-15B 高出 21 - 44 个百分点,置信区间不重叠。在 arith+func 和模板化参数化的 StableHLO-Held-Out-200 上,验证器语义取决于属性值而非结构,相同的基线匹配或超过了 SLM,将这些情况视为非优势情况。最后发布了基准测试、解码器、所有提示生成以及一个可重现的 Docker 镜像。

英文摘要

Multi-Level Intermediate Representation (MLIR) underlies modern ML compiler infrastructure (TensorFlow, JAX/StableHLO, PyTorch Inductor, IREE), yet appears only in trace amounts in code-LM pretraining corpora. MLIR is also extensible by design: new dialects ship per application domain, so a fine-tuned model per dialect does not scale. We ask whether inference-time priors derived mechanically from each dialect's Operation Definition Specification (ODS) can substitute for gradient-based adaptation. First, we release four natural-language-to-MLIR benchmarks across three dialects - MLIR-Spec-150, Linalg-Spec-30, StableHLO-Spec-30, and StableHLO-Held-Out-200 - totaling 410 in-scope NL-to-MLIR pairs, plus a 25-program out-of-grammar stress set and a hand-authored n=30 functional reference set, shipped under Apache-2.0 with Gebru datasheets and Croissant 1.0 metadata. Second, we build a three-layer schema-derived constraint stack: a CFG over op signatures(C1), type-domain splits from an ODS-extracted type lattice (C2), and an SSA-scope validator driving five-retry rejection sampling (C3). Porting from arith+func+memref+linalg to StableHLO required no new constraint-layer code. On dialects whose verifier semantics are dominated by structural constraints, schema-derived priors let SmolLM2-1.7B match or exceed 15B-34B code LMs at 8-25x the per-generation speed: on linalg, SmolLM2 reaches 80.0% verify-valid (three-seed mean, n=125), beating CodeLlama-34B, Granite-Code-34B, and StarCoder2-15B by 21-44 percentage points with non-overlapping CIs. On arith+func and on the templated parametric StableHLO-Held-Out-200, where verifier semantics turn on attribute values rather than structure, the same baselines match or beat the SLM; we scope these as non-win cells. We release benchmarks, decoder, all per-prompt generations, and a reproducibility Docker image.

发表机构

  • Meta

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑