arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33411cs.AIcs.CLcs.SE

MetaBench-Harness:解锁基准测试框架的端到端优化

MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses

Xuanjun Chen, Hua-Hsuan Chen, Wei-Chung Lu, Yinghao Ma, Jyh-Shing Roger Jang, Hung-yi Lee

AI总结:

针对静态基准快速饱和的问题,提出MetaBench-Harness双循环搜索框架,端到端优化基准生成流程,在CodeContests和AIME-2024上生成对前沿模型有挑战性和区分度的基准,实现多维度进化。

AI中文摘要:

大型语言模型(LLM)的快速发展使得静态基准测试的饱和速度超过了其设计速度。现有的自动进化框架试图通过扰动单个任务来生成更难的问题,但它们仍受限于僵化、硬编码的生成规则。超越孤立任务的进化,我们提出用MetaBench-Harness对基准生成工作流本身进行端到端优化,这是一个双循环搜索框架。具体而言,内循环利用基准测试框架在每轮中生成一个新的基准,而外层的元框架编排层则基于历史进化轨迹,对框架实现进行迭代式细化和搜索。通过将MetaBench-Harness应用于编程竞赛数据集CodeContests和奥林匹克数学数据集AIME-2024,我们证明了进化出的基准对前沿模型具有挑战性和区分度。轨迹与质量分析验证了MetaBench-Harness能够实现多维度进化,在连续轮次中稳步提升进化的合理性、基准的胜任能力以及评估器的鲁棒性。此外,案例研究揭示了其有效利用多种难度杠杆来重构问题并提升所需能力。最终,这项工作为基准饱和这一紧迫挑战提供了解决方案。

英文摘要:

Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid, hard-coded generation rules. Moving beyond the evolution of isolated tasks, we propose to optimize the benchmark generation workflow itself end to end with MetaBench-Harness, a dual-loop search framework. Specifically, the inner loop utilizes a benchmark harness to generate a new benchmark in each round, while the outer meta-harness orchestration layer iteratively refines and searches over harness implementations based on historical evolution trajectories. By applying MetaBench-Harness to the competitive programming CodeContests and Olympiad mathematics AIME-2024 datasets, we demonstrate that the evolved benchmarks are challenging and discriminative for frontier models. Trajectory and quality analyses verify that MetaBench-Harness enables multi-dimensional evolution, steadily improving evolution reasonableness, benchmark competency, and evaluator robustness across successive rounds. Furthermore, case studies reveal its effective utilization of diverse difficulty levers to reframe problems and elevate required capabilities. Ultimately, this work provides a solution to the pressing challenge of benchmark saturation.

补充信息

↑