arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Learn2Zinc:微调小型语言模型以实现MiniZinc中的文本到模型翻译

Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc

Serdar Kadioglu, Karthik Uppuluri

arXiv 2607.20456首次发表:更新:

发表机构

AI Center of Excellence, Fidelity Investments; Department of Computer Science, Brown University(富达投资卓越人工智能中心; 布朗大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对小型语言模型在MiniZinc文本到模型翻译的难题,提出跨模型错误引导方法,收集语法错误构建训练数据集微调模型,提升执行准确率至98%,虽语法可学但约束推理仍具挑战,且开源相关资源供后续研究。

AI 中文摘要

大型语言模型在主流编程语言的代码生成方面表现出色,但在处理诸如MiniZinc(一种用于组合问题的约束建模语言)等罕见的领域特定语言时存在困难。我们研究了有针对性的微调是否能让小型语言模型(参数从0.6B到20B)从自然语言问题描述中生成语法正确且语义有效的MiniZinc模型。我们发现语法错误在处理该领域特定语言时占主导地位,如Qwen3、LLaMa、Gemma和GPT - OSS等小型语言模型的开箱即用执行准确率接近零。我们提出了一种跨模型错误引导方法,收集多个大语言模型运行中的语法错误并利用它们策划一个纠错训练数据集。通过该数据集微调小型语言模型,能持续提高所有模型规模下的直接代码生成和思维链方法。通过自我反思和集成,我们的方法实现了高达98%的执行准确率。同时,解决方案准确率仍为35%,这表明虽然语法是可学习的,但约束推理仍然是一个挑战。我们将微调管道、数据集和模型开源,以供文本到模型翻译的进一步研究。

英文摘要

Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such as MiniZinc, a constraint modeling language for combinatorial problems. We investigate whether targeted fine-tuning can teach small language models (0.6B to 20B parameters) to generate syntactically correct and semantically valid MiniZinc models from natural language problem descriptions. Our key finding is that syntax errors dominate failures when working with this domain specific language: the out-of-the-box execution accuracy of small language models such as Qwen3, LLaMa, Gemma, and GPT-OSS is near-zero. We propose a cross-model error bootstrapping approach that collects syntax errors from multiple LLM runs and leverage those to curate an error correction training dataset. This dataset allows us fine-tune small language models that consistently improves both direct code generation and chain-of-thought approaches across all model sizes. With self-reflection and ensembling, our approach achieves up to 98\% execution accuracy. In parallel, solution accuracy still remains at 35\%, indicating that while syntax is learnable, constraint reasoning remains a challenge. We contribute our fine-tuning pipeline, datasets, and models to opens-source for further research on text-to-model translation.

CommentsCP 2026 Workshop on LLMs meet Constraint Solving

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑