arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越准确率:大型语言模型统计推理的多维度评估

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

Monnie McGee, Mateo Langston Smith, Julian Cabrera

arXiv 2608.03038首次发表:更新:

AI 中文总结

本研究结合多维度分析框架评估15款LLMs的统计推理,发现仅用准确率无法全面描述其统计推理能力,且存在厂商专属解释风格差异。

AI 中文摘要

统计推理具有多维度特征,但大型语言模型(LLMs)的评估通常侧重响应准确率,却忽略了模型构建与传达统计解释的方式。本研究结合响应准确率、响应行为、结构主题建模与词汇相似度分析,展示了多维度评估的价值。该框架应用于15款当前主流LLMs针对90道题目生成的解释,这些题目来自高中、本科及研究生阶段的4套统计学考试。模型间准确率差异显著,介于55%至78%之间;结构主题建模显示,所有模型的统计推理存在共同概念组织,而词汇相似度分析则发现了适度但一致的厂商专属解释风格差异——同一厂商(如Anthropic、OpenAI)开发的模型生成的解释,比不同厂商模型的解释略为相似。这些结果表明,当代LLMs的统计推理不能仅用准确率来描述,且对响应行为与模型生成解释的补充分析,可为生成式AI的统计推理提供更全面的评估。

英文摘要

Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how models construct and communicate statistical explanations. This study demonstrates the value of a multidimensional evaluation by combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis. The framework is applied to explanations generated by 15 current-generation LLMs responding to 90 questions drawn from four statistics examinations spanning high school, undergraduate, and graduate levels. Accuracy varied substantially across models, ranging from 55\% to 78\%. In contrast, structural topic modeling revealed a common conceptual organization of statistical reasoning across all models, while lexical similarity analysis identified modest but consistent vendor-specific differences in explanatory style. Models developed by the same vendor (e.g. Anthropic, OpenAI) produced explanations that were slightly more similar than models from different vendors. These findings demonstrate that statistical reasoning in contemporary LLMs cannot be characterized by accuracy alone and illustrate how complementary analyses of response behavior and model-generated explanations provide a more comprehensive evaluation of statistical reasoning in generative AI.

Comments15 pages, 5 tables, 2 figures, presented at JSM 2026 and submitted for publication

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑