arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlavourBench:基于可执行烹饪基准真值的前沿语言模型排名

FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training

Josef Chen, Erim Hayretci

arXiv 2608.20574首次发表:更新:

发表机构

Imperial College London(帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FlavourBench是一款自动化语言模型排名基准,以可执行烹饪系统为基准真值,评估27个前沿语言模型,发现Grok 4.6表现最优,可有效消除排行榜差异缺失,结果可靠。

AI 中文摘要

开放式语言模型基准通常需要一个评判者:人类偏好小组、另一个模型,或脆弱的精确匹配密钥。我们推出FlavourBench,这是一个自动化基准,其中版本化烹饪系统提供密集的可执行基准真值。每个任务提供8种食材,要求生成3种食材的组合;在模型执行前,Epicure会对所有56种可能的组合进行评分。我们在相同的534个核心任务(涵盖替代、配对和受限组合)上评估27个前沿端点。每个排名模型在每个小组和系列中恰好有89个有效响应(总计14418个模型-任务单元),消除了排行榜的差异缺失。FlavourBench分数是冻结任务分数的同系列平均值。我们使用50000个锚定集群自举重复来生成同时95%分数带,并使用100000个符号翻转抽取来进行所有351个配对模型对比,采用Holm控制。两个独立编制的小组的相关性为r=0.89(秩相关rho=0.80)。Grok 4.6的点估计值最大,为65.1(同时95%置信区间61.0-69.2);351个模型对中有101个得到区分。发布内容包括提示、所有组合分数图、原始响应、精确路径、内容哈希以及可重建所有结果的离线验证器。

英文摘要

We introduce FlavorBench: a benchmark for Compiling Dense Deterministic Answer Maps from a Versioned Culinary Embeddings Model. We test 27 frontier large language model endpoints on 534 substitution, pairing and constraining tasks for tasks that request a 3-ingredient portfolio from 8 candidates and score all 56 resulting portfolios. We conducted multiplicity-controlled paired tests on 101 of 351 model contrasts for this task-set. The largest point estimate on this task-set was achieved by Grok 4.6 at 65.1. The same rankings for this task-set were also achieved on several independently-compiled panels (using a variety of familiar metrics, task filters, etc.) and 3 public Epicure checkpoints. We present a 3-seed post-training study where LoRA SFT of a Qwen3-0.6B checkpoint on 270 optimal answers for Epicure to score on this task-set resulted in a 13.3 point gain on 84 anchor-disjoint maps (compared to format and label-matched control; 95% CI: 6.52, 20.29; p = 0.000170).

Comments18 pages, 11 figures. Evaluation of 27 frontier language-model endpoints on 534 identical tasks per model, comprising 14,418 scored model-task cells. Adds reward-map sensitivity, selection and metric robustness, held-out Recipe1MSubs substitution validation, and a preregistered controlled reward-transfer study. Code, dataset, and interactive leaderboard links remain unchanged

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑