arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过训练进行编译:将自然语言规范转换为局部神经函数

Compile by Training: Turning Natural-Language Specifications into Local Neural Functions

Yuntian Deng, Pengyu Nie, Stuart Shieber

arXiv 2609.04199首次发表:更新:

发表机构

University of Waterloo; Harvard University(滑铁卢大学; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出“通过训练进行编译”方法,将自然语言规范转为可复用神经函数,在FuzzyBench-Hard上达83.6%语义准确率,可部署于多场景。

AI 中文摘要

许多重复的文本功能易于描述,但难以通过规则实现;而对每个输入都调用大型远程模型会带来重复成本、延迟以及对提供商的依赖。我们提出了“通过训练进行编译”的方法,该方法可将自然语言规范转换为可复用的神经函数。在编译阶段,教师模型会生成特定任务的示例,这些示例用于训练紧凑解释器的小型适配器。生成的函数无需教师模型即可运行,并且可以像普通软件一样存储、版本控制和组合。在FuzzyBench-Hard(Program-as-Weights快速编译器未产生任何精确匹配的子集)上,“通过训练进行编译”达到了83.6%的语义准确率。这种更高的准确率伴随着更高的编译成本:大约一分钟,而快速编译器仅需几秒。我们将该编译器部署在公共交互式服务中,并在多站点网站助手、语言控制的3D化身以及英语-Claudish双向翻译器中展示了编译后的函数。

英文摘要

Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At compile time, teacher models generate task-specific examples that are used to train a small adapter for a compact interpreter. The resulting function runs without the teachers and can be stored, versioned, and composed like ordinary software. On FuzzyBench-Hard, a subset on which the Program-as-Weights fast compiler produced no exact matches, compile by training reaches 83.6% semantic accuracy. This higher accuracy comes with a higher compile-time cost: roughly a minute rather than seconds for the fast compiler. We deploy the compiler in a public interactive service and demonstrate compiled functions in a multi-site website helper, a language-controlled 3D avatar, and a bidirectional English-Claudish translator.

CommentsEMNLP 2026 System Demonstrations. Demo: https://programasweights.com

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑