arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06933cs.CLcs.AI

Ask-E:一个用于校准型问题生成的环境

Ask-E: An Environment for Calibrated Question Generation

发表机构华盛顿大学 · 艾伦人工智能研究所
查看机构详情
  • University of Washington(华盛顿大学)
  • Allen Institute for AI(艾伦人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

Sarah Pratt, Jae Sung Park, Scott Geng, Ali Farhadi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出Ask-E环境,以两个现有语言模型的能力范围定义目标技能水平,通过生成仅能被其中一个模型解决的问题来校准模型,该环境可用于基准测试与训练,且训练后模型在下游数学基准上表现提升。

中文摘要 AI 辅助

如今,我们通过在模型能力前沿的问题上进行训练和评估来改进模型。创建此类问题本身就是一项艰巨的任务,需要具备探测模型极限、泛化至现有问题分布之外的能力,还需要将问题设置在精确的难度级别,这要求理解解决这些问题所需的条件。简而言之,生成校准至模型当前前沿的问题需要具备超出该前沿的能力,随着模型性能提升,这一约束会愈发繁重。我们的核心见解是,可将这一约束转化为优势:能够持续生成校准至给定前沿的问题的模型,必然具备超出该前沿的能力。因此,我们提出Ask-E,这是一个针对模型编写给定技能水平问题的能力进行基准测试和训练的环境,而非针对回答问题的能力。具体而言,我们将目标技能水平定义为两个现有语言模型能力所界定的范围。若生成的问题恰好能被这两个模型中的一个解决,则该问题成功实现校准,精确落在目标范围内,且能区分这两个模型的能力。Ask-E兼具基准测试和训练环境的功能,模型可在此生成校准至不同技能水平的问题。我们发现,即使是前沿模型在该基准上的校准准确率也低于50%,为衡量未来进展留下了巨大空间。我们还表明,在该环境中进行训练,即使不使用新的数学数据、不与更强模型交互、也不使用基于正确性的奖励,也能在多个下游数学基准上实现性能提升。

英文摘要

Today, we improve models by training and evaluating them on problems at the frontier of their abilities. Creating such problems is itself a demanding task, requiring the ability to probe model limits and generalize beyond existing question distributions. It also means placing problems at a precise difficulty level, which requires understanding what it takes to solve them. In short, generating problems calibrated to a model's current frontier demands capability beyond it, an increasingly burdensome constraint as models improve. Our key insight is that we can leverage this constraint to our advantage: a model that can generate problems consistently calibrated to a given frontier must possess capability beyond it. Accordingly, we present Ask-E, an environment that benchmarks and trains models on their ability to write questions at a given skill level, rather than answer them. Concretely, we define target skill levels as ranges bounded by the capabilities of two existing language models. A generated question is successfully calibrated if exactly one of the two models can solve it, placing it precisely within the target range and differentiating the capabilities of these models. Ask-E serves both as a benchmark and a training environment, where models generate problems calibrated to a variety of skill levels. We find that even frontier models achieve below 50% calibration on the benchmark, leaving significant headroom to measure future progress. We also show that training on this environment leads to improvements across a number of downstream math benchmarks even with no new math data, no interaction with stronger models, and no correctness-based reward.

↑