arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22067cs.CLcs.AI

基于美国核管理委员会反应堆操作员执照考试的多模态语言模型微调与检索策略基准测试

Multimodal Language Models Benchmarked Against the NRC Reactor Operator Licensing Examination: Fine-Tuning and Retrieval Strategies

Isak Hwang, Yoon Pyo Lee, Syed Bahauddin Alam

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对美国核管理委员会反应堆操作员执照考试,评估310亿参数多模态模型应用核知识的能力,通过对比基础模型与多种微调及检索配置,发现固定大小分块RAG的SFT配置表现最佳,并揭示了分块策略规律及RAFT与SFT的性能差异。

中文摘要 AI 辅助

将大语言模型集成到核电行业需要基于特定领域知识的输出。本研究通过针对美国核管理委员会反应堆操作员执照考试对八种模型检索配置进行基准测试,评估了一个拥有310亿参数的开放权重多模态模型(Gemma 4 31B-IT)应用核知识的能力。使用标准的80%人类通过标准,评估了2015年至2021年3月场次的14项通用基础知识考试(7项压水堆和7项沸水堆考试)。将基础模型与利用Gemini提炼的思维链(CoT)原理进行监督微调(SFT)、对美国能源部基础手册使用BM25稀疏检索的检索增强生成(RAG)以及检索增强微调(RAFT)的配置进行比较。在检索管道内,比较了固定大小滑动窗口分块与结构感知分块。带有固定大小分块RAG的SFT配置在14项考试中的8项中达到了标准,优于所有其他配置,而没有微调的配置没有一项通过。总体准确率达到79.7%,置信区间跨越阈值,压水堆项目的准确率为80.2%。此外,还出现了两个规律:首选的分块策略根据模型的训练状态而反转,并且在匹配搜索环境中,RAFT的表现不如标准SFT。这些结果表明了微调与搜索方法的哪种组合能够实现操作员级别的能力。

英文摘要

Competence claims for a language model in a safety-critical domain are credible when measured against a standard the domain already enforces. We evaluate an open-weight 31-billion-parameter multimodal model (Gemma 4 31B-IT) on the U.S. Nuclear Regulatory Commission Reactor Operator Generic Fundamentals Examination (GFE), scoring it paper by paper against the 80% criterion applied to every human candidate, with no rounding up. The evaluation set is a census of every GFE administered at the March sitting from 2015 to 2021, giving seven pressurized water reactor (PWR) and seven boiling water reactor (BWR) papers and 697 scored items. Eight configurations cross three model states, the base model, supervised fine-tuning (SFT) on distilled chain-of-thought rationales and retrieval-augmented fine-tuning (RAFT), with three retrieval conditions, none and BM25 retrieval over the Department of Energy Fundamentals Handbooks under fixed-size and structure-aware chunking. Out of the box it answers 51.94% correctly and passes no paper. SFT with fixed-size chunking retrieval passes 8 of 14, reaching 80.23% on PWR items and 79.77% pooled, with a Wilson interval spanning the threshold. The preferred chunking granularity reverses with training state, structure-aware before fine-tuning and fixed-size after, so chunking optimized against a base model cannot be inherited by its fine-tuned descendant. RAFT trails SFT by 2.2 to 2.3 percentage points overall, and the deficit holds in all four reactor-type and chunking strata. The pipeline runs on one workstation with no network access at run time, and the result approaches operator-level command of engineering fundamentals without reliably achieving it.

发表机构

  • organization= Department of Nuclear Engineering, Hanyang University , addressline= 222 Wangsimni-ro , postcode= 04763 , state= Seongdong-gu , city= Seoul , country= South Korea
  • organization= The Grainger College of Engineering, Nuclear, Plasma \& Radiological Engineering, University of Illinois Urbana-Champaign , city= Urbana , state= IL , country= USA

机构由 AI 辅助整理,请以论文原文为准。

↑