发表机构
Blue Whale Lab, National University of Singapore; Hong Kong Polytechnic University; Fudan University(新加坡国立大学蓝鲸实验室; 香港理工大学; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究化学推理语言模型中思维链的作用,发现其在多个模型和任务中易产生幻觉且与答案正确性无关,存在共享草稿本功能,如Chem-R依赖SMILES草稿,干扰其草图会影响生成,表明化学CoT是易产生幻觉的分子草稿本,提醒不能仅以CoT判断推理可靠性。
AI 中文摘要
化学推理语言模型期望通过可靠的思维链(CoT)得出分子答案。然而,在四个推理模型家族和十二项化学任务中,幻觉普遍存在,且在很大程度上与答案正确性无关:正确答案常与相关分子中不存在的虚构结构断言共存。归因分析表明,以特定于模型的形式存在共享的草稿本功能:Chem-R和ether-0依赖于碎片化的SMILES草稿,而ChemDFM-R强调支架、位置和命名线索。值得注意的是,干扰Chem-R的SMILES草图会降低生成效果,这表明即使语言结构断言大多无作用时,结构草稿也可能具有因果承载作用。这些结果表明,化学CoT既不是可靠的解释,也不仅仅是事后的合理化,而是一个易产生幻觉的分子草稿本。这一发现提醒人们不要将CoT视为可靠推理的直接证据,并促使人们在仅评估答案之外进行过程级监督。
英文摘要
Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning model families and twelve chemistry tasks, hallucination is widespread and largely decoupled from answer correctness: correct answers often coexist with fabricated structural claims absent from the relevant molecules. Yet this does not make the reasoning trace computationally irrelevant. Attribution analyses suggest a shared scratchpad function expressed in model-specific forms: Chem-R and ether-0 rely on fragmented SMILES drafts, whereas ChemDFM-R emphasizes scaffold, positional, and naming cues. Notably, perturbing Chem-R's SMILES sketches degrades generation, showing that structural drafts can be causally load-bearing even when verbal structural claims are largely inert. Together, these results show that chemical CoT is neither a faithful explanation nor merely a post-hoc rationalization, but a hallucination-prone molecular scratchpad. This finding cautions against treating CoT as direct evidence of faithful reasoning and motivates process-level supervision beyond answer-only evaluation.
Comments17 pages, 6 figures