arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21504cs.LG

ChemDIRT:面向稳健化学大语言模型评估的多样化指令、表示与任务基准

ChemDIRT: A Diversified Instruction, Representation, and Task Benchmark for Robust Chemistry-LLM Evaluation

Eric Inae, Tim Gunn, Chris Bond, Meng Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对现有化学LLM评估基准的局限性,提出ChemDIRT基准,涵盖八大类化学任务,评估模型在指令和分子表示变化下的准确性与一致性,测试发现LLM存在提示敏感性、表示依赖及任务间性能不均问题。

中文摘要 AI 辅助

大型语言模型(LLM)的快速发展推动了其在化学等科学领域的应用,但现有化学基准往往仅能提供模型能力的狭窄视角,聚焦于有限的任务集,却忽略了模型对问题表述和化学表示变化的稳健性,导致报告的性能可能高估了模型在现实场景中进行一致推理的真实能力。为应对这一挑战,我们提出ChemDIRT(Diversified Instruction, Representation, and Task Benchmark,即多样化指令、表示与任务基准),这是一个用于评估LLM化学推理稳健性的综合评估框架。ChemDIRT系统地测量模型在指令和分子表示变化下的性能,涵盖八大类化学任务,通过在这些受控扰动下同时评估准确性和一致性,相比传统的单一格式基准,ChemDIRT能更可靠地评估模型的推理能力。我们对一系列开源和闭源LLM进行了基准测试,结果显示这些模型存在显著的提示敏感性、表示依赖性以及不同任务类别间的性能差异。

英文摘要

The rapid advancement of large language models (LLMs) has led to increasing interest in their application to scientific domains such as chemistry. However, existing chemistry benchmarks often provide only a narrow view of model capability, focusing on limited task sets while overlooking robustness to variations in problem formulation and chemical representation. As a result, reported performance may overestimate a model's true ability to reason consistently across realistic settings. To address this challenge, we introduce ChemDIRT (Diversified Instruction, Representation, and Task Benchmark), a comprehensive evaluation framework designed to assess the robustness of chemical reasoning in LLMs. ChemDIRT systematically measures model performance across variations in instructions and molecular representations while spanning eight categories of chemistry tasks. By evaluating both accuracy and consistency under these controlled perturbations, ChemDIRT provides a more reliable assessment of model reasoning capabilities than conventional single-format benchmarks. We benchmark a diverse set of open- and closed-source LLMs, revealing substantial prompt sensitivity, representation dependence, and uneven performance across task families.

发表机构

  • University of Notre Dame(圣母大学)

机构由 AI 辅助整理,请以论文原文为准。

↑