arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00070cond-mat.mtrl-scics.AI

AutoXRD:用于粉末衍射分析的自主大语言模型智能体及综合评估

AutoXRD: Autonomous LLM Agents and Comprehensive Evaluation for Powder Diffraction Analysis

发表机构香港中文大学 · 香港理工大学
查看机构详情
  • The Chinese University of Hong Kong(香港中文大学)
  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuetong Wu, Maojun Sun

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AutoXRD自主LLM智能体框架,构建含双赛道的XRDBench评估10种LLM,发现其在XRD分析关键环节仍存不足,验证了AutoXRD组件的有效性并明确了改进方向。

中文摘要 AI 辅助

粉末X射线衍射(XRD)是材料表征的核心手段,但可靠的端到端自动化仍具挑战性。XRD智能体需解读衍射证据、操作精修软件、按合理顺序管理耦合参数,并区分数值改进与物理有效性。本文提出AutoXRD,这是一个自主大语言模型(LLM)智能体框架,将粉末XRD分析组织为逐步精修过程,使动作基于观测证据,并在接受结果前应用确定性晶体学与物理检查。我们进一步引入含两个互补赛道的XRDBench:XRDBench-QA包含100个受限诊断任务,用于分离科学推理与决策;XRDBench-E2E包含34个可执行工作流,测试智能体能否将这些能力组合为完整分析,需完成文件检查、晶体学软件执行、迭代精修、证据保留及报告。我们在1340次模型-任务运行中评估了10种近期LLM,模型平均得分仅57.8(满分100),从XRDBench-QA的61.9降至XRDBench-E2E的53.7;它们在精修历史评估与结果接受上表现最佳,但在精修动作选择、相量化、索引及Rietveld精修上仍明显较弱。GPT-5.6 Sol获最高综合得分81.1,GPT-5.6 Terra获最高XRDBench-E2E点估计值81.0,GPT-5.6 Luna获最佳得分-成本权衡。消融实验显示AutoXRD的全部6个组件均持续提升性能,支持该框架设计。最后,执行轨迹分析揭示了耦合参数控制、定量推理、证据保留及工作流终止方面的反复失败,为更强的科学约束、不确定性感知决策及更高效规划提供了方向。

英文摘要

Powder X-ray diffraction (XRD) is central to materials characterization, yet reliable end-to-end automation remains challenging. An XRD agent must interpret diffraction evidence, operate refinement software, manage coupled parameters in a defensible order, and distinguish numerical improvement from physical validity. In this paper, we propose AutoXRD, an autonomous large language model (LLM) agent framework that organizes powder-XRD analysis as stepwise refinement, grounds actions in observed evidence, and applies deterministic crystallographic and physical checks before accepting results. We further introduce XRDBench with two complementary tracks. XRDBench-QA contains 100 bounded diagnostic tasks that isolate scientific reasoning and decision-making, whereas XRDBench-E2E contains 34 executable workflows that test whether agents can compose these capabilities into complete analyses requiring file inspection, crystallographic-software execution, iterative refinement, evidence preservation, and reporting. We evaluate ten recent LLMs across 1,340 model--task runs. Models average only 57.8 out of 100, falling from 61.9 on XRDBench-QA to 53.7 on XRDBench-E2E. They perform best on refinement-history assessment and result acceptance, but remain substantially weaker on refinement-action selection, phase quantification, indexing, and Rietveld refinement. GPT-5.6 Sol achieves the highest overall score of 81.1, GPT-5.6 Terra the highest XRDBench-E2E point estimate of 81.0, and GPT-5.6 Luna the best score--cost trade-off. Ablations show that all six AutoXRD components consistently improve performance, supporting the framework design. Finally, execution-trace analysis reveals recurring failures in coupled-parameter control, quantitative reasoning, evidence preservation, and workflow termination, motivating stronger scientific constraints, uncertainty-aware decisions, and more efficient planning.

↑