CARAT:材料大语言模型是在推理还是背诵?
CARAT: Do Materials LLMs Reason or Recite?
浏览论文内容
中文总结 AI 辅助
CARAT通过八种匹配视图和消融实验证明,材料大语言模型在结构问答中常背诵而非推理,并引入可弃权规则与监督来揭示和缓解该问题。
中文摘要 AI 辅助
当材料大语言模型回答关于晶体结构的问题时,它是在基于结构进行推理,还是复制其输入中已经打印的答案?准确率无法区分:结构描述往往打印出与评分目标完全相同的字段。CARAT在八个匹配视图下保持问题和金标准答案不变,在GraphSpace中分别命名每个结构关系,并加入匹配微调、答案掩蔽、证据注入、配对推理以及一个可以弃权(不执行)的规则。首先,在基准测试中最难的家族上,基于grounded视图相比公式输入高出17.3个百分点。其次,我们将这种审视转向自身。GraphSpace比普通周期图高出19.3个百分点,但这一差距同时包含两种效应:在普通渲染携带问题所需一切信息的情况下,差距为1.96个百分点;而在其完全省略这些字段的情况下,差距为46.7个百分点。头条数字主要衡量的是基线所缺乏的内容,而非证据的呈现方式。第三,我们攻击自己的基准。一个跳过链接直接读取列表的规则在七个加固家族中答对了四个,因此我们重建了它,直到十一个此类捷径接近随机水平。冻结模型引用该链接,但当我们重定向它时,在95.6%的配对案例中回答相同:它重复了关系却没有使用它。经过匹配监督后,这一比例达到99.8%,而删除链接则使其降至23.4%,低于最佳捷径达到的27.0%:这两个步骤都是可学习的。
英文摘要
When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accuracy cannot tell: a structural description often prints the very field it is scored against. CARAT holds question and gold answer fixed across eight matched views, names each structural relation separately in GraphSpace, and adds matched fine-tuning, answer masking, evidence injection, paired inference, and a rule that can withhold claims. First, on the benchmark's hardest families the grounded view is worth 17.3 points over formula inputs. Second, we turn that scrutiny on ourselves. GraphSpace beats a plain periodic graph by 19.3 points, but that margin is two effects at once: where the plain rendering carries everything the question needs it is 1.96 points, and where it omits those fields entirely, 46.7 points. The headline mostly measures what the baseline lacked, not how evidence is presented. Third, we attack our own benchmark. A rule that skips the link and reads the list directly answers four of seven hardened families, so we rebuilt it until eleven such shortcuts sat near chance. The frozen model quotes that link yet answers the same when we redirect it, on 95.6% of paired cases: it repeats the relation without using it. After matched supervision it reaches 99.8%, and deleting the link drops it to 23.4%, below the 27.0% the best shortcut reaches: both steps are learnable.
发表机构
- Beihang University(北京航空航天大学)
- Imperial College London(伦敦帝国理工学院)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。