FindStatBench:在组合代码合成上评估大语言模型
FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis
浏览论文内容
中文总结 AI 辅助
介绍用于评估大语言模型组合代码合成能力的FindStatBench基准,涵盖多种合成任务,通过特定方式评分并评估11个系统,发现开源闭源系统准确率趋同、示例可能有害、存在输出预算问题等,揭示统计合成与映射合成的差异及相关问题。
中文摘要 AI 辅助
我们介绍了FindStatBench,这是一个用于在组合代码合成上评估大语言模型的执行基准。它基于FindStat构建,包含24个集合中的2329个任务和552万个隐藏实例,涵盖统计合成(将对象映射到整数)和映射合成(将对象映射到对象)。每个任务给出数学描述和至多五个公共输入输出示例,模型必须在无检索、工具使用等情况下输出一个Python求解函数。通过对保留的组合对象进行精确的沙盒执行来评分。我们评估了11个系统,发现了三个主要模式:最强的开源和闭源系统在实例准确率上相差1个百分点内;示例可能有害;一些失败反映了输出预算机制。总体而言,统计合成比映射合成容易得多,一些集合的准确率仍接近零,长提示会导致准确率急剧下降,精确的符号规则归纳仍然很脆弱。
英文摘要
We introduce FindStatBench, an execution benchmark for evaluating large language models on combinatorial code synthesis. Built from FindStat, it contains 2,329 tasks across 24 collections and 5.52M hidden instances, covering statistic synthesis, which maps objects to integers, and map synthesis, which maps objects to objects. Each task gives a mathematical description and at most five public input-output examples; a model must emit one Python solve function with no retrieval, tool use, execution feedback, voting, or reranking. Submissions are scored by exact sandboxed execution on held-out combinatorial objects. We evaluate eleven systems: four closed-source production models and seven open-weight models served through one inference provider. FindStatBench reveals three main patterns. First, the strongest open- and closed-source systems converge within 1 pp instance accuracy, and both an oracle over all systems and five-way sampling from one mid-tier model yield only limited task-accuracy gains. Second, examples can hurt: several classical bijections are solved perfectly with zero examples but fail under five-example prompts. Third, some failures reflect output-budget mechanics, as reasoning can exhaust the visible response before code is emitted. Overall, statistic synthesis is much easier than map synthesis, some collections remain near-zero, long prompts cause a sharp accuracy cliff, and exact symbolic rule induction remains brittle.