arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从表格到量化陈述:通过可执行验证评估LLM推理生成

From Tables to Quantified Statements: Evaluating LLM Inference Generation through Executable Verification

Mai Mohamed Eida, Gunjan Anand, Ayush Singh, Aleksandre Maskharashvili

arXiv 2609.23966首次发表:更新:

发表机构

University of Illinois Urbana Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出STAT-TO-TEXT任务,利用LLM从统计表格生成量化推理,并通过可执行Python代码验证,发现更大模型(如GPT-OSS-120B)在忠实度、覆盖率和多样性上表现更优。

AI 中文摘要

大型语言模型(LLM)能够从表格生成流畅的描述,但其输出可能在逻辑上不受结构化数据的支持。我们引入了STAT-TO-TEXT,一个受控任务,在该任务中,LLM使用诸如all、some、no和most等量化结构,从统计表格生成量化的自然语言推理。为了评估这些推理,我们使用LLM生成的Python检查器代码,该代码在执行时根据表格验证相应的真值条件。我们比较了四个开放权重LLM,涵盖不同模型系列和规模,评估忠实度、逻辑准确性、表格覆盖率和多样性。我们的结果表明,模型规模和系列很重要,最大的模型(GPT-OSS-120B)始终产生最忠实的推理,同时不牺牲更大的表格覆盖率和量词多样性,而较小的模型则不然。这些发现得到了人工标注的支持,表明自动化检查器与人类判断高度一致。

英文摘要

LLMs can generate fluent descriptions from tables, but their outputs may remain logically unsupported by the structured data. We introduce STAT-TO-TEXT, a controlled task in which LLMs generate quantified natural language inferences from statistical tables using quantified constructions such as all, some, no, and most. To evaluate these inferences, we use an LLM generated Python checker code which when executed verifies the corresponding truth conditions against the table. We compare four open-weight LLMs across model families and scales, evaluating faithfulness, logical accuracy, table coverage, and diversity. Our results show that model scale and family matter, with the largest model (GPT-OSS-120B) consistently producing the most faithful inferences without sacrificing greater table coverage and quantifier diversity, as opposed to smaller models. These findings are supported by human annotation, which shows that the automated checker closely aligns with human judgments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑