arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

微调小型语言模型以生成可靠的VASP INCAR文件

Fine-Tuning Small Language Models for Reliable VASP INCAR Generation

Xinyue Zhang, Jixiang Li, Bin Shao, Baishun Yang, Zhiyang Liu, Weichao Wang

arXiv 2608.05387首次发表:更新:

AI 中文总结

本研究通过微调小型语言模型(SLM)并搭配后处理器VASPGuard构建INCAR-SLM,在INCARBench上优于通用大模型,证明小型模型经适配可可靠生成VASP INCAR文件,适配后性能随规模增长趋于饱和。

AI 中文摘要

语言模型可根据自然语言请求生成VASP INCAR文件,但目前仅大型专有云模型能可靠处理紧密耦合、对物理参数敏感的设置,这种依赖与本地高通量材料工作流(注重隐私、成本和离线部署)不匹配。本文表明小型语言模型(SLM)可缩小这一差距:该SLM基于参考VASP计算进行微调,并搭配确定性后处理器VASPGuard(用于检查语法、工作流及材料相关约束),组合模型名为INCAR-SLM。在VASP INCAR生成基准INCARBench上,基于Qwen3-4B的INCAR-SLM优于所有评估的通用大语言模型,在100分制INCAR得分上超过GPT-5.4达15.55分。该提升主要来自微调,VASPGuard用于修正剩余错误。进一步发现模型规模影响低于预期:应用微调与后处理后,性能在数十亿参数时趋于饱和,Qwen3-4B优于同系列更大模型。

英文摘要

Language models can prepare VASP INCAR files from natural-language requests, but so far only large proprietary cloud models come close to handling the tightly coupled, physics-sensitive settings reliably, a dependence that fits poorly with local, high-throughput materials workflows where privacy, cost, and offline deployment matter. We show that a small language model (SLM) can close this gap. The SLM is fine-tuned on reference VASP calculations and paired with VASPGuard, a deterministic post-processor that checks syntax, workflow, and material-dependent constraints; we call the combined model INCAR-SLM. On INCARBench, a benchmark for VASP INCAR generation, INCAR-SLM built on Qwen3-4B outperforms every general-purpose LLM evaluated, exceeding GPT-5.4 by 15.55 points on the 100-point INCAR Score. Most of this gain comes from fine-tuning, with VASPGuard correcting the errors that remain. We further find that model size matters less than expected: once fine-tuning and post-processing are applied, performance saturates at a few billion parameters, and Qwen3-4B outperforms larger models in the same family.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑