arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在开放权重语言模型中读取和引导材料科学机制的表示

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

Markus J. Buehler

arXiv 2607.20058首次发表:更新:

发表机构

Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型中材料科学机制信息的表示形式,结合多种方法,包括匹配读数、状态几何等,通过实验验证其三种形式,如概念可读、取向由状态变换承载等,还通过比较提示等发现物理关系在受控状态变化中更易显现。

AI 中文摘要

大语言模型能回答科学问题,但正确输出不表明其是否体现或运用了主导物理原理。研究表明开放权重google/gemma - 4 - E4B - it模型中的材料科学机制信息有三种可通过实验分离的形式:概念在单个隐藏状态中可读,本构取向由状态间的受控变换承载,选定的内部表示因果性地控制工程答案。研究结合了匹配的直接和雅可比词汇读数、无选项状态几何、60定律反事实基准和因果干预。在保留的50个材料描述中得到了相关结果,还通过比较特定提示等方法进一步验证,发现物理关系在受控状态变化中比仅在绝对状态中更明显。

英文摘要

Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here, using three open-weight Gemma 4 models (google/gemma-4-E4B-it, google/gemma-4-12B-it, google/gemma-4-31B-it) we identify three experimentally separable signatures of materials-science mechanism information: selective concept readability, relational encoding of qualitative constitutive orientation, and causal, context-dependent control of constrained engineering answers. We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions. In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blinded identification of 9 of 10 mechanism families. A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison. We therefore compared otherwise identical prompts in which only the direction of the physical input was reversed, asking whether the resulting hidden-state movement followed the supplied constitutive law. These state transformations ordered direct, physically neutral and inverse laws across 60 frozen relations and correctly oriented 39 of 40 directional laws, whereas lexical controls were near chance. Bidirectional interventions shifted answer probabilities toward or away from the physically appropriate outcome across all 12 matched cases, while counterfactual state patches transferred opposing decision signals across mechanisms and answer formats. Physical relationships were therefore more visible in controlled state changes than in absolute states alone.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑