arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SleuthBench:使用表格隐藏信号对统计LLM评估进行基准测试

SleuthBench: Benchmarking Statistical LLM Evaluation Using Tabular Hidden Signals

Jingyun Jia, Antoine Remond-Tiedrez, Aaron Alvarez, Joshua Shunk, Rich Caruana, Ben Lengerich

arXiv 2609.34228首次发表:更新:

发表机构

University of Wisconsin–Madison; Intelligible; University of Cincinnati(威斯康星大学麦迪逊分校; Intelligible; 辛辛那提大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SLEUTHBENCH通过注入可控模式到公共表格数据中,自动生成基准真相,评估LLM智能体在数据质量与特征贡献上的统计发现能力,并引入经验层预计算工件将特征贡献准确率从41.9%提升至68.0%。

AI 中文摘要

使用大型语言模型(LLM)智能体进行统计发现需要可验证的分析基准真相。为真实世界数据集建立这样的基准真相成本高昂,而且对公共数据集的先验知识可能影响智能体的响应。我们引入了SLEUTHBENCH,一个通过将受控的数据质量问题和特征效应注入公共表格数据集来解决这两个问题的基准:注入的模式决定了答案,因此参考答案自动计算,对原始表格的记忆知识不足,而表格保持其背景结构。注入的模式基于真实数据分析中报告的现象进行建模。该基准定义了17个问题模板,分为两个系列:数据质量问题和特征贡献问题。我们使用Python编码工具评估了六个最先进的LLM,对70个经过验证的数据集-模板组合的数据科学和商业表述进行分析,总共产生1680个分级响应。模型可靠地检测数据质量问题(准确率83.8%),但恢复特征贡献的能力较差(41.9%)。发现特征如何塑造目标需要在候选变量和分析程序之间进行搜索。为解决这个问题,我们提出了经验层,一组预计算的统计工件,包括摘要、拟合的特征和交互效应以及数据集描述,这些工件暴露了候选模式以供直接检查。访问这些工件将特征贡献准确率从41.9%提高到68.0%。

英文摘要

Evaluating statistical discovery by large language model (LLM) agents requires verifiable analytical ground truth. Establishing such ground truth for real-world datasets is costly, and prior knowledge of public datasets can influence agent responses. We introduce SLEUTHBENCH, a benchmark that addresses both problems by injecting controlled data-quality problems and feature effects into public tabular datasets: the injected pattern determines the answer, so reference answers are computed automatically and memorized knowledge of the original table is insufficient, while the table keeps its background structure. The injected patterns are modeled on phenomena reported in real data analyses. The benchmark defines 17 question templates in two families: data-quality questions and feature-contribution questions. We evaluate six state-of-the-art LLMs that analyze the data using a Python coding tool, on data-science and business phrasings of 70 validated dataset-template combinations, yielding 1680 graded responses in total. The models detect data-quality problems reliably (83.8% accuracy) but recover feature contributions poorly (41.9%). Finding how features shape the target requires searching over both candidate variables and analytical procedures. To address this issue, we propose the Empirical Layer, a set of precomputed statistical artifacts comprising summaries, fitted feature and interaction effects, and dataset descriptions, which exposes candidate patterns for direct inspection. Access to these artifacts raises feature-contribution accuracy from 41.9% to 68.0%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑