arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用国际象棋基准测试大语言模型的提示优化

Benchmarking Prompt Optimization of Large Language Models With Chess

Timothée Lesort, Alejandra López de Aberasturi Gómez, Tristan Karch, Tom Veniat, Philippe Modard, Karl Tuyls, Ludovic Denoyer

arXiv 2610.00416首次发表:更新:

发表机构

imec AI.labs(imec AI实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型自动提示优化评估难的问题,提出基于1118个Lichess谜题的国际象棋基准,验证六种APO算法在八个模型上的效果,具有挑战性、区分性、可更新且成本约800美元。

AI 中文摘要

随着大语言模型能力的提升,对其进行评估变得越来越具有挑战性:基准测试可能饱和,公开测试集存在污染风险,而评估更难的任务可能需要昂贵的评分或执行基础设施。这些挑战在自动提示优化(APO)中被放大,因为在搜索更好提示的过程中会重复进行评估。因此,研究APO需要一个评分便宜且确定性高、难度足以留有改进空间、并且随着模型发展可更新的基准。我们引入了一个基于1,118个Lichess谜题构建的国际象棋基准,用于研究冻结大语言模型的APO:我们优化其提示而不更新模型权重。国际象棋结合了廉价的精确匹配评分、基于引擎的替代走法评估,以及难度可调且可更新的问题供应。与仅报告孤立测试项成功率的评估不同,该基准还将谜题解决收益与同一领域内的短对局模拟联系起来。我们用它来评估八个目标模型上的六种APO算法,不仅测量基线强度,还测量每个模型对优化的响应程度,以及优化后的提示是否跨模型迁移并应用于对局。因此,国际象棋非常适合作为APO的基准:它(i)具有挑战性,因为即使是最强的评估模型Gemini 3.5 Flash(用作元模型)也只能解决约55%的谜题;(ii)具有区分性,能揭示不同方法和模型之间的收益、无变化和性能回退;(iii)可更新,提供新谜题以减少污染风险,并具有可调难度以在模型改进时保持提升空间;以及(iv)成本低廉,整个研究运行约需800美元。我们发布了谜题、优化和评估代码以及数据集更新脚本(此HTTPS URL)。

英文摘要

Evaluating large language models becomes increasingly challenging as their capabilities advance: benchmarks can saturate, public test sets risk contamination, and assessing harder tasks can require expensive grading or execution infrastructure. These challenges are amplified in automatic prompt optimization (APO), where evaluation is repeated throughout the search for better prompts. Studying APO therefore requires a benchmark that is cheap and deterministic to score, hard enough to leave room for improvement, and renewable as models evolve. We introduce a chess benchmark built from 1,118 Lichess puzzles to study APO for frozen LLMs: we optimize their prompts without updating their model weights. Chess combines inexpensive exact-match scoring, engine-based evaluation of alternative moves, and a renewable supply of problems with adjustable difficulty. Unlike evaluations that report only success on isolated test items, the benchmark also connects puzzle-solving gains to short game-play rollouts within the same domain. We use it to evaluate six APO algorithms on eight target models, measuring not only baseline strength but also how much each model responds to optimization and whether optimized prompts transfer across models and to game play. Chess is thus a well-suited benchmark for APO: it is (i) challenging, as even the strongest evaluated model, Gemini 3.5 Flash (used as the meta-model), solves only about 55\% of puzzles; (ii) discriminative, revealing gains, unchanged performance, and regressions across methods and models; (iii) renewable, with fresh puzzles to reduce contamination risk and adjustable difficulty to maintain headroom as models improve; and (iv) affordable, as the complete study runs for around \$800. We release the puzzles, optimization and evaluation code, and dataset-renewal scripts (https://github.com/imec-ailabs/Automatic-Prompt-Optimization-with-Chess).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑