发表机构
Washington University in St. Louis; Oak Ridge National Laboratory; Lawrence Berkeley National Laboratory; UniverseTBD; Massachusetts Institute of Technology; Old Dominion University(圣路易斯华盛顿大学; 橡树岭国家实验室; 劳伦斯伯克利国家实验室; UniverseTBD; 麻省理工学院; 奥多明尼昂大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对适配的 Qwen3-8B 模型 Graph-PRefLexOR-8B,构建可视化诊断工作流以追踪材料科学假说生成的图到答案机制,发现机制恢复集中于后期合成及答案起始层,可助力相关人员识别假说的机制支持变化。
AI 中文摘要
AI 协作科学家可生成流畅的材料科学假说,但流畅性不代表答案保留了具有科学意义的机制。我们针对 Graph-PRefLexOR-8B(一款为暴露头脑风暴、图构建、模式提取及合成等不同阶段而调整的 Qwen3-8B 模型)开展了从图到答案的机制追踪案例研究。我们将语义回溯、图破坏、基于激活的恢复测量及按 token 区域分层网格整合为可视化诊断工作流,以检查该路径。在 100 个开放式材料科学问题中,最终答案最接近模型自身的结构化阶段,尤其是合成阶段。在图破坏条件下,对 37 个残差流检查点(含嵌入输出及 36 个 transformer 块)的全面扫描显示,在第 7-10 层的早期过渡区域几乎无机制恢复,恢复反而集中在第 30 和 36 层附近的后期合成及答案起始区域。该工作流旨在帮助科学家和模型开发者在生成的假说传递给下游实验规划前,识别其何时失去或重获机制支持。
英文摘要
AI co-scientists can generate fluent materials-science hypotheses, but fluency does not show that an answer preserves a scientifically meaningful mechanism. We present a graph-to-answer mechanism-tracing case study for Graph-PRefLexOR-8B, a Qwen3-8B model adapted to expose distinct stages for brainstorming, graph construction, pattern extraction, and synthesis. We organize semantic backtracking, graph corruption, activation-based recovery measurements, and layer-by-token-region grids into a visual diagnostic workflow for inspecting this pathway. Across 100 open-ended materials-science questions, final answers remain closest to the model's own structured stages, especially synthesis. Under graph corruption, a full sweep over 37 residual-stream checkpoints, the embedding output and 36 transformer blocks, shows little mechanism recovery in the earlier transition region at layers 7--10, recovery instead concentrates in late synthesis and answer-start regions around layers 30 and 36. The workflow is intended to help scientists and model developers identify where a generated hypothesis loses or regains mechanism support before it is passed to downstream experimental planning.