化学与材料推理的有根据验证:检测是瓶颈
Grounded verification of chemical and materials reasoning: detection is the bottleneck
浏览论文内容
中文总结 AI 辅助
研究针对大语言模型在化学推理中虚构对象的问题,提出分层验证器,通过与数据库核对及门控校正减少公式错误,检索次数大幅减少,修复成功率高,但检测召回率是瓶颈,基于事实验证可提高答案质量,长尾错误处提升明显。
中文摘要 AI 辅助
大语言模型在流畅的推理过程中会虚构化学对象(分子式、空间群、形成能),尤其集中在置信度最不可靠的长尾实体上。确定性的、基于数据库的验证可以捕捉并修复此类错误,且无需全面检索的覆盖成本。研究发现,约束环节在于检测而非修复。分层验证器提取每个可检查的断言,与权威数据库和物理学进行核对,并将参考信息输入门控校正循环。在四个模型和528个条件固定的提示中,门控校正将公式错误率从22%降至4%,检索次数比全面扩充少$3.2\times$,超过了对话预言机。修复在标记触发时成功率较高(80 - 97%),瓶颈在于循环检测召回率。只有当验证器的范围达到可交付成果时,基于事实的验证才会提高最终答案的质量(从83%提高到90%),且提升仅出现在存在可提取长尾错误的地方:在接近上限的物理常数上没有提升,在同位素半衰期上提升较大(从11%降至0%)。
英文摘要
Language models are moving into chemistry and materials discovery workflows, where a wrong molecular formula, space group, or formation energy can silently propagate into downstream decisions. These confabulations hide inside fluent reasoning traces and concentrate on rare, long-tail entities, where model confidence is least trustworthy. Retrieving reference data for every prompt would catch them, but at a heavy coverage and abstention cost. We show that deterministic, database-grounded verification catches and repairs these errors selectively, and that the binding constraint is detection rather than repair. Our tiered verifier extracts each checkable claim, tests it against authoritative databases and physical law, and retrieves a reference value only when a check fails. Across four models and over five hundred prompts with pinned conditions, gated correction cuts the error rate of committed formulas from 22% to 4% with 3.2 times fewer retrievals than blanket augmentation, and it outperforms a conversational retrieval oracle when every answer, corrected or not, is scored. When a flag fires, repair almost always succeeds; the benefit reaches the final answer only where the verifier's scope covers it and where long-tail error exists. Checkable claims, checked cheaply, are a practical lever for trustworthy machine reasoning in chemistry.