AI 中文总结
研究围绕科学发现评估大语言模型的能力,引入SDABench基准,涵盖多领域多能力的真实和合成数据实例,评估15个模型,发现其在描述性分析好但在特定任务下降,还提供错误分析框架定位失败之处。
AI 中文摘要
现有的科学数据分析基准主要从代码执行或工作流程完成方面评估大语言模型,忽视了科学分析旨在支持不同类型的科学主张,如假设探索、统计推断、机制解释等,各有不同假设和有效性标准。我们引入了SDABench基准,它围绕五个领域(生物、化学、环境、地理、物理)的六种能力(描述性、探索性、推断性、预测性、因果性和机制性)重新组织评估。SDABench包含527个真实数据实例和6000个合成实例,有选择题和开放式两种格式。评估15个代表性大语言模型发现,模型在描述性分析方面表现良好,但在需要假设选择、潜在过程建模或机制推理的任务上急剧下降。SDABench还提供了一个五阶段错误分析框架来定位大语言模型的失败之处。
英文摘要
Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, each with different assumptions and validity criteria. We introduce SDABench, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics). SDABench comprises 527 real-data instances (SDA-Real) and 6000 synthetic instances (SDA-Synth), each in both multiple-choice and open-ended formats, constructed through an automated pipeline. Evaluating 15 representative LLMs, we find that models handle descriptive analysis well but degrade sharply on tasks requiring assumption selection, latent-process modeling, or mechanistic reasoning. SDABench further provides a five-stage error analysis framework that locates where LLMs fail: more advanced models more reliably identify the relevant scope and variables, but still struggle to select appropriate analytical procedures, model variable relationships, and draw valid conclusions.