arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估用于智能体DFT工作流的4B开放权重本地大语言模型:文献可复现性审计

Evaluating a 4B open-weights local LLM for agentic DFT workflows: a literature reproducibility audit

Shambhu Bhandari Sharma

arXiv 2608.29665首次发表:更新:

发表机构

University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究审计4B开放权重本地大语言模型Qwen3:4B在智能体DFT工作流中的表现,发现其参数提取精度较高,完整GPU驻留提升性能,可重现部分文献的晶格常数,证明轻量开放权重模型可可靠驱动自主智能体工作流。

AI 中文摘要

材料科学中依赖托管商业模型的智能体工作流面临严重的可复现性、经济性和数据隐私约束。为探索完全本地的智能体科学,本研究评估了开放权重模型Qwen3:4B在不同硬件约束下执行自主科学流程的表现。该系统应用于五角二维材料,从非结构化文本中提取参数,将其转化为密度泛函理论(DFT)输入,并在严格的神经符号架构下驱动模拟收敛,其中智能体提出方案,确定性代码执行处理。该工作流通过逐字文本 grounding 和多轮推理组合来应对硬件诱导的结构崩溃。针对201项专家判断进行评估,提取器的精确率达到95.7%(95%置信区间90.3-98.1%),召回率为67.3%(59.8-74.0%),确保提取的参数严格符合事实。然而,识别缺失参数的精确率未超过47.0%,表明测量的遗漏率是文献真实不完整性的宽松上限。在三种硬件配置下,完整GPU驻留对提取质量的影响比权重或缓存精度更根本,在固定量化水平下使马修斯相关系数从0.414提升至0.530,使用未量化缓存时则提升至0.560。语料库规模审计显示,57项研究中仅有19项(33.3%)原则上可复现,即报告了重新初始化计算所需的全部方法参数。在驱动至收敛后,该工作流重现了已发表的晶格常数,平均绝对相对误差为2.3%,且弛豫结构保留其原型,证明轻量级开放权重模型在确定性代码门限约束下可可靠驱动自主智能体工作流。

英文摘要

Agentic workflows in materials science relying on hosted commercial models face severe reproducibility, economic, and data-privacy constraints. To explore fully local agentic science, this work evaluates an open-weights Qwen3:4B model executing an autonomous scientific pipeline across varying hardware constraints. Applied to pentagonal two-dimensional materials, the system extracts parameters from unstructured text, translates them into density functional theory (DFT) inputs, and drives simulations to convergence under a strict neurosymbolic architecture where agents propose and deterministic code disposes. The workflow is guarded by verbatim text grounding and multi-pass inference unions to counteract hardware-induced structural collapse. Evaluated against 201 expert judgements, the extractor achieves 95.7% precision (95% CI 90.3-98.1%) and 67.3% recall (59.8-74.0%), ensuring extracted parameters are strictly factual. However, precision identifying absent parameters does not exceed 47.0%, establishing that the measured omission rate constitutes a loose upper bound on true literature incompleteness. Across three hardware configurations, complete GPU residency governs extraction quality more fundamentally than weight or cache precision, raising Matthews correlation from 0.414 to 0.530 at fixed quantisation and to 0.560 with an unquantised cache. A corpus-scale audit indicates only 19 (33.3%) of the 57 studies are reproducible in principle, reporting every method parameter needed to re-initialise the calculation. Driven to convergence, the workflow reproduces published lattice constants with a mean absolute relative error of 2.3% where the relaxed structure retains its prototype, establishing that lightweight open-weights models can reliably drive autonomous agentic workflows when bounded by deterministic code gates.

Comments14 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑