OenoBench:用于大型语言模型知识基础评估的葡萄酒领域基准
OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models
查看机构详情
- DipWSET(葡萄酒与烈酒教育基金会文凭)
- StrategAI
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究推出葡萄酒领域基准OenoBench,构建含3266道题的语料库,评估16种LLM配置,发现不同模型准确率差异及推理模式提升等规律并发布相关资源。
中文摘要 AI 辅助
我们推出OenoBench,这是一个葡萄酒领域知识基准,包含3266道多项选择题,涵盖六个支柱(产区、葡萄品种、葡萄栽培、酿酒、生产商、商业)和四个难度等级。该语料库由35个经来源验证的爬虫从政府登记处(INAO、TTB、OIV)、同行评审期刊以及维基百科/维基数据中提取的38104个原子级、来源锚定的事实构建而成。我们的方法学贡献是一个由大语言模型(LLM)驱动的流程,其中语言模型对经验证的事实进行格式化并审核结果,但绝不作为事实的来源:每一项主张都可追溯到一个URL,每一个问题都由五个生成器类别中的五种策略之一生成,并且每一个问题都由九个智能体组成的审核组根据人类黄金标准通过Cohen's κ进行校准。通过评估16种前沿配置,我们发现:(i)整体准确率在53%至84%之间,o3以83.6%领先;(ii)推理模式提升集中在DeepSeek R1(+6.8个百分点),而Claude Opus和Gemini Pro无此提升;(iii)Anthropic在其自身问题上表现出+9个百分点的自我偏好,而Google则表现出-8个百分点的反向偏好;(iv)前沿开放权重模型与专有推理模型共享成本-准确率帕累托前沿;(v)每种配置在闭卷可解答项上均提升约33个百分点,揭示了参数化召回上限,仅上下文切片可避免该上限。我们根据CC-BY-SA-4.0协议发布语料库、审核结果、人工审核应用程序及构建代码。
英文摘要
We introduce OenoBench, a wine-domain knowledge benchmark of 3,266 multiple-choice questions across six pillars (regions, grape varieties, viticulture, winemaking, producers, business) and four difficulty tiers. The corpus is built from 38,104 atomic, source-anchored facts extracted by 35 provenance-verified scrapers from government registries (INAO, TTB, OIV), peer-reviewed journals, and Wikipedia/Wikidata. Our methodological contribution is an LLM-driven pipeline in which language models reformat verified facts and audit the result, but never serve as the source of truth: every claim traces to a URL, every question is generated by one of five strategies across five generator families, and every question is scored by a nine-agent audit calibrated against a human gold sheet via Cohen's $κ$. Evaluating sixteen frontier configurations, we find: (i) overall accuracy spans 53%-84%, led by o3 at 83.6%; (ii) reasoning-mode lift concentrates in DeepSeek R1 (+6.8pp) and is absent in Claude Opus and Gemini Pro; (iii) Anthropic shows +9pp self preference on its own questions while Google shows -8pp inverse preference; (iv) frontier open-weight models share the cost-vs-accuracy Pareto frontier with proprietary reasoning models; and (v) every config gains around 33pp on closed-book solvable items, revealing a parametric-recall ceiling that only the contextual slice avoids. We release corpus, audit findings, human-review app, and construction code under CC-BY-SA-4.0.