发表机构
Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对统计问题形式化任务构建基准StatFormBench,评估14个LLMs在该任务的两个子任务上的表现,发现现有模型性能有限且无一致最优者,仅提示策略提升效果受限。
AI 中文摘要
大型语言模型(LLMs)正越来越多地被用作统计与数据科学工作的助手,但现有评估大多假设分析目标已被明确指定。在实际应用中,用户带着非正式的目标和异构数据而来,需要模型来确定隐含的统计任务以及哪些数据相关。我们首先将这一上游步骤正式定义为统计问题形式化(Statistical Problem Formulation),并将其分解为两个子任务:(1)统计问题分类,(2)变量识别与角色分配。随后我们推出StatFormBench,一个由五本跨领域统计教科书和一个数据科学案例库构建的基准,涵盖多样的问题类型、数据表示和场景风格,包含1013个样本,覆盖20个粗粒度和85个细粒度统计问题类别。在14个开源与闭源LLMs中,表现最佳的零样本模型仅达到72.0的细粒度分类准确率和63.2的变量集重叠度;没有模型在两个子任务上均表现最优,而增强的提示策略仅产生有限或不一致的提升。我们在Hugging Face发布了基准数据,在GitHub发布了评估代码。
英文摘要
Large language models (LLMs) are increasingly used as assistants for statistical and data science work, yet existing evaluations largely assume the analysis target is already specified. In practice, users arrive with informal goals and heterogeneous data, leaving the model to decide what statistical task is implied and which data are relevant. We first formalize this upstream step as Statistical Problem Formulation and decompose it into two subtasks: (1) Statistical Problem Classification and (2) Variable Identification & Role Assignment. We then introduce StatFormBench, a benchmark built from five cross-domain statistics textbooks and a data science case library, covering diverse problem types, data representations, and scenario styles. It contains 1,013 samples spanning 20 coarse-grained and 85 fine-grained statistical problem categories. Across 14 open- and closed-source LLMs, the best zero-shot models reach only 72.0 fine-grained classification accuracy and 63.2 variable set overlap. No model performs consistently best across the two subtasks, while enhanced prompting strategies yield only limited or inconsistent gains. We release the benchmark data on Hugging Face at https://huggingface.co/datasets/THU-CongLab/StatFormBench and the evaluation code on GitHub at https://github.com/THU-CongLab/StatFormBench.
CommentsAccepted for publication at the EMNLP 2026 main conference