arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Harness修复了什么?一项关于可见性、基线充分性与评估缺陷的预注册研究

What Does a Harness Repair? A Preregistered Study of Visibility, Baseline Adequacy and Evaluation Defects

Bowen Xu, Boyu Chen

arXiv 2610.05533首次发表:更新:

发表机构

Stanford University; University College London(斯坦福大学; 伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过预注册实验探究Harness搜索增益来源,发现关闭思考设置常优于受限思考,GEPA修复起点但无更优,评估缺陷可被复制而非揭示,强调基线与评估可见性的重要性。

AI 中文摘要

Harness搜索会保留对冻结模型周围的提示、推理开关、令牌预算或解析器的更改,如果该更改提高了分数。这种增益可能来自解析器之前无法读取的答案、薄弱的比较或评估中的缺陷。我们预先注册了一项研究,探讨这些增益的来源,涉及三个小模型、三个基准、复制和测试分区、一个GEPA搜索臂以及六种一次注入一个的评估缺陷,并报告了全部47个主要终点。在9个模型-基准组合中,关闭思考在5个组合中相对于受限思考设置提高了准确率,且在每种情况下,增益主要来自受限设置无法给出可读答案的问题。关闭思考设置在13项比较中的11项中并未显著差于救援配置或四个GEPA选择的harness,并在GSM8K上输给了两个模型的救援配置。GEPA修复了其损坏的起点,但其选择的harness中没有一个比关闭思考设置更准确。服务引擎中的思考预算(也允许更长的答案)在9个组合中的6个中降低了截断并提高了解析率。在15个可评估的缺陷-模型对中,有6个通过相同管道的复制再现了缺陷的扭曲而非揭示它。在LongevityBench多项选择任务中,只有经过长寿调优的模型击败了最强的恒定标签基线。

英文摘要

Harness search keeps a change to the prompts, reasoning switches, token budgets or parsers around a frozen model if the change raises a score. Such a gain can come from answers the parser could not read before, a weak comparison, or a defect in the evaluation. We preregistered a study of where these gains come from, with three small models, three benchmarks, replication and test partitions, a GEPA search arm and six evaluation defects injected one at a time, and we report all 47 primary endpoints. Turning thinking off raised accuracy over a capped thinking setting in 5 of 9 model-benchmark cells, and in each the gain came mostly from questions where the capped setting gave no readable answer. The thinking-off setting was not meaningfully worse than a rescue configuration or four GEPA-selected harnesses in 11 of 13 comparisons, and lost to the rescue on GSM8K for two models. GEPA repaired its broken starting points, but none of its selected harnesses was more accurate than the thinking-off setting. A thinking budget in the serving engine, which also allows a longer answer, lowered truncation and raised the parse rate in 6 of 9 cells. In 6 of 15 evaluable defect-model pairs, replication through the same pipeline reproduced the defect's distortion instead of revealing it. On the LongevityBench multiple-choice tasks, only the longevity-tuned model beat the strongest constant-label baseline.

Comments43 pages, 2 figures, 35 tables. Ancillary files in anc/: the frozen preregistration, its addenda (row keys of five rows withheld) and the aggregate analysis report, with a README

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑