发表机构
University of Michigan; Michigan Institute for Computational Discovery and Engineering (MICDE)(密歇根大学; 密歇根计算发现与工程研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过微调开源扩散模型构建核能文生成像模型,发现微调效果依赖生成架构,且微调后开源模型在专业核能图像生成上优于主流商业系统,证实领域特定微调是开发可信领域生成式AI的可行路径。
AI 中文摘要
生成式人工智能(AI)已实现文本到图像合成的变革,但其对专业工程领域的表征能力仍未得到充分探索。以核工程为例,通用基础模型常生成物理层面错误或概念不一致的图像,原因在于其缺乏领域特定知识。本研究针对核能领域的文生成像任务,开展了首批针对开源扩散模型的领域适配系统性研究。我们整理了包含反应堆、燃料循环、辐射及相关概念的1000张带标注核能图像数据集,并用其微调了三款最先进的开源模型:Stable Diffusion XL(SDXL)、SD-v3.5-Medium及基于流匹配的Flux.1模型。我们采用定量图像相似度指标与定性专家评估,将微调后模型与对应零样本模型对比评估性能。微调显著提升了SDXL的保真度,仅为SD-v3.5-Medium带来有限增益,且未对Flux.1产生可测提升,表明适配效果强烈依赖底层生成架构,而非仅模型规模。我们还将微调后模型与三款领先商业系统——GPT-Image-2、Gemini-3.1-Flash-Image及Midjourney对比。尽管GPT-Image-2与Gemini能为宽泛核能概念生成可信图像,但常无法满足专业工程提示词的要求,而微调后的开源模型能产出更准确、技术上更一致的结果。这些结果证实,领域特定微调是开发适用于专业领域的可信生成式AI工具的可行路径。
英文摘要
Generative artificial intelligence (AI) has transformed text-to-image synthesis, yet its ability to represent specialized engineering domains remains largely unexplored. As an exmaple in nuclear engineering, general-purpose foundation models frequently generate physically incorrect or conceptually inconsistent images because they lack domain-specific knowledge. This work presents one of the first systematic studies of domain adaptation for nuclear text-to-image generation through fine-tuning of open-source diffusion models. We curate a dataset of 1,000 captioned nuclear energy images spanning reactors, fuel cycles, radiation, and related concepts, and use it to fine-tune three state-of-the-art open-source models: Stable Diffusion XL (SDXL), SD-v3.5-Medium, and the flow-matching Flux.1 model. Their performance is evaluated using both quantitative image-similarity metrics and qualitative expert assessment against the corresponding zero-shot models. Fine-tuning substantially improves the fidelity of SDXL, provides only limited gains for SD-v3.5-Medium, and yields no measurable improvement for Flux.1, demonstrating that adaptation effectiveness depends strongly on the underlying generative architecture rather than model scale alone. We further compare the fine-tuned models against three leading commercial systems--GPT-Image-2, Gemini-3.1-Flash-Image, and Midjourney. Although GPT-Image-2 and Gemini generate convincing images for broad nuclear concepts, they frequently fail on specialized engineering prompts, where the fine-tuned open-source models produce more accurate and technically consistent outputs. These results establish domain-specific fine-tuning as a practical pathway for developing trustworthy generative AI tools for domain-specific applications.
Comments29 pages, 10 figures, and 4 tables