发表机构
The Chinese University of Hong Kong; Zhongguancun Academy; SINOPEC Research Institute of Petroleum Processing; Beijing Normal-Hong Kong Baptist University; Guangzhou Nanfang College(香港中文大学; 中关村学院; 中国石化石油化工科学研究院; 北京师范大学-香港浸会大学联合国际学院; 广州南方学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过受控实验分离覆盖率与专门化,发现LLM工具框架的重复执行掩盖了专门化,提出任务优势需跨执行持久且可指导决策的评估标准。
AI 中文摘要
自动生成LLM工具框架有望通过任务专门化来改进推理。然而,额外的答案覆盖率可能来自同一程序的重复执行,这使得专门化难以识别。我们引入了一种受控评估,将答案覆盖率、可重复的任务优势以及执行前选择带来的增益分离开来。在386个MATH-500任务上,我们比较了八个生成的工具框架加上一个基线,其中基线有九个字节相同的副本,每个成员执行三次。相同的程序产生了2.16个百分点的重复平均预言机余量。生成的程序表现出明显更多的可重复得分模式,但这些主要揭示了持续的弱点:在100个任务上,相对于基线的损失在所有三次重复中持续存在,而持续的优势仅出现在一个任务上,并且对答案提取敏感。冻结的选择器获得了0.00个百分点的增益,两个群体在27次工具框架执行时达到了98.70%的预言机覆盖率。稳定的互补性在三次重复时仍未解决。支持性的BIRD轨迹定位了机制实现、激活和输出有效性方面的失败。总之,这些发现确立了为什么仅凭覆盖率和可重复性不能证明有用专门化的主张。它们为工具框架多样性提出了一个评估标准:任务优势应在执行中持续存在,指导可用的决策,并在匹配的推理预算下改进额外的固定程序执行。
英文摘要
Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advantages, and gains from pre-execution selection. On 386 MATH-500 tasks, we compare eight generated harnesses plus a baseline with nine byte-identical baseline copies, using three executions per member. Identical programs yield 2.16 percentage points of repeat-averaged oracle headroom. Generated programs exhibit substantially more repeatable score patterns, but these chiefly reveal persistent weaknesses: losses relative to the baseline persist across all three repeats on 100 tasks, while persistent wins occur on only one task and are sensitive to answer extraction. The frozen selector gains 0.00 percentage points, and both populations reach 98.70% oracle coverage at 27 harness executions. Stable complementarity remains unresolved at three repeats. Supporting BIRD traces locate failures in mechanism implementation, activation, and output validity. Together, these findings establish why coverage and repeatability alone cannot justify claims of useful specialization. They motivate an evaluation standard for harness diversity: task advantages should persist across executions, guide usable decisions, and improve on additional fixed-program executions under matched inference budgets.
Comments26 pages, 4 figures, including references and appendices. Under review at ICLR 2027. Code: https://github.com/StatXzy7/harness-eval