发表机构
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对粒子物理中FORM语言缺乏AI辅助的问题,提出验证驱动数据生成流水线并微调Qwen3-8B模型,在四个基准上超越千亿参数前沿模型,同时保持通用能力。
AI 中文摘要
FORM是一种领域特定的符号操作语言,广泛用于粒子物理中处理多圈费曼图计算产生的大型代数表达式。尽管它在精密理论物理中扮演核心角色,但据我们所知,目前尚无人工智能工具辅助物理学家编写FORM代码。我们证明,当代大型语言模型(LLMs),包括具有数千亿参数的前沿模型,在没有文档的单次尝试中,在我们的指令遵循和教程式FORM任务上实现了零百分比的执行通过率,这确立了FORM在撰写本文时对LLMs而言是一种真正的零样本语言。接着,我们提出了一种验证驱动的数据生成流水线,该流水线使用FORM二进制本身作为执行预言机,生成并验证了包含4,633个训练示例的语料库,涵盖确定性计算、开放式程序、教程代码和知识问答对。使用量化低秩适配(QLoRA)微调一个紧凑的开源权重模型(Qwen3-8B),产生了一个专家模型,在四个互补基准(840个任务,每个任务单次尝试)上的评估显示,在执行率和较大基准上的严格FORM验证输出匹配方面,它显著优于参数高达756B的前沿模型,并且在较小但更难的基准上与它们统计上无法区分。通用推理和编码能力保持在2.6个百分点以内。
英文摘要
FORM is a domain-specific symbolic manipulation language widely used in particle physics for processing the very large algebraic expressions arising from multi-loop Feynman diagram calculations. Despite its central role in precision theoretical physics, no artificial-intelligence tooling exists, to our knowledge, for assisting physicists in writing FORM code. We show that contemporary large language models (LLMs), including frontier models with hundreds of billions of parameters, achieve a zero-percent execution pass rate on our instruction-following and tutorial-style FORM tasks without documentation in a single attempt, establishing FORM as a genuine zero-shot language for LLMs at the time of writing. We then present a verification-driven data generation pipeline that uses the FORM binary itself as an execution oracle to produce and validate a corpus of 4,633 training examples spanning deterministic computations, open-ended programs, tutorial code, and knowledge question-answer pairs. Fine-tuning a compact open-weights model (Qwen3-8B) with quantized low-rank adaptation (QLoRA) yields a specialist that, evaluated on four complementary benchmarks (840 tasks, single attempt each), decisively outperforms frontier models with up to 756B parameters in execution rate and in strict, FORM-verified output matching on the larger benchmarks, and remains statistically indistinguishable from them on the smaller, harder ones. General reasoning and coding capabilities are preserved within 2.6 percentage points.