发表机构
Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出COGTRL这一轨迹级强化学习框架,训练LLMs产生科学发现所需的认知轨迹,在AI和材料科学领域的实验中,其方法质量较3B参数模型基线平均提升7.85分,表现接近70B参数模型,且更受领域专家偏好。
AI 中文摘要
在大量科学研究上训练的大语言模型(LLMs)正越来越多地被整合为科学发现的助手。然而,大多数研究论文省略了实现预期目标所需的检查约束、失败的替代方案和迭代决策的细粒度认知过程。这类认知过程对在约束下朝着特定目标工作的现实世界科学家至关重要。在本文中,我们表明,当LLMs被训练以产生此类认知轨迹时,它们作为科学发现助手的表现要优于仅在科学文献上训练的情况。我们提出了COGTRL,这是一种轨迹级强化学习框架,通过联合优化认知轨迹和以交错方式产生的科学步骤,训练LLMs模拟基于认知的推理。在两个3B参数模型和两个科学领域(AI与材料科学)中,COGTRL相较于可比的3B参数模型基线,方法质量平均提高了7.85分,且达到了与70B参数模型相当的性能。此外,领域专家的分析显示,与基线相比,他们更偏好COGTRL生成的方法。
英文摘要
Large Language Models (LLMs) trained on extensive scientific research are increasingly integrated as assistants for scientific discovery. However, most research papers omit the fine-grained cognitive process of examining constraints, failed alternatives, and iterative decisions required to achieve the desired goal. Such cognitive processes are vital for real-world scientists working toward specific goals under constraints. In this paper, we show that LLMs, when trained to produce such cognitive traces, perform better as scientific discovery assistants than when trained solely on scientific literature. We propose COGTRL, a trajectory-level reinforcement learning framework that trains LLMs to emulate cognitively grounded reasoning by jointly optimizing cognitive traces and the scientific steps produced in an interleaved manner. Across two 3B-parameter models and two scientific domains (AI and Materials Science), COGTRL improves method quality by an average of 7.85 points over comparable 3B model baselines and achieves competitive performance relative to 70B parameter models. Moreover, analysis by domain experts shows a preference for methods generated by COGTRL over the baselines.
CommentsAccepted EMNLP 2026