什么让 Terminal-Bench 任务困难?在裁决型智能体语料库上区分真实难度与虚假难度
What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus
AI总结:
本研究分析Terminal-Bench 3/Frontier-Bench 0.1的生产记录,提出区分真实难度与虚假难度的方法,发现仅125个全失败任务中78个为认证未解决,强调基准应报告全失败任务证据。
AI中文摘要:
前沿基准需要当前模型无法解决的任务。但一个没有模型能解决的任务并不自动就是困难任务。相同的零通过率可能源于真实的能力差距,但也可能源于缺失上下文、参考解决方案损坏、基础设施故障或可被绕过的验证器。在本文中,我们使用包含1,081个拉取请求、639个评分任务、28,801次试验和105,933美元记录的智能体花费的冻结Terminal-Bench 3 / Frontier-Bench 0.1生产记录来研究这一问题。我们探究一个全失败任务实际上证明了什么。对于125个无诚实通过的任务,我们结合任务工件、参考解决方案运行、空解决方案对照、对抗性试验、轨迹、遥测和审查记录,并应用有序有效性筛选。125个任务中只有78个作为认证未解决候选存活。其余任务包括14个具有损坏预言机、8个受基础设施故障主导、4个仅能通过验证器绕过通过,以及21个其可解性未被现有证据认证。因此,缺乏饱和与真正困难不是同一回事。认证未解决标签也是狭窄的:它意味着作者路线通过、基础设施未主导、未观察到严格绕过,且所有评估代理均失败。它并不证明内在难度、验证器完整性或预期能力上的失败。我们进一步分析被拒绝的提交和通过的任务,以表明仅通过率无法解释任务为何困难。总体而言,我们的结果表明,前沿基准在将全失败任务用作能力声明之前,应报告其背后的证据。
英文摘要:
Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come from a real capability gap, but it can also come from missing context, a broken reference solution, infrastructure failure, or a verifier that can be bypassed. In this paper, we study this issue using a frozen Terminal-Bench 3 / Frontier-Bench 0.1 production record with 1,081 pull requests, 639 scored tasks, 28,801 trials, and $105,933 in logged agent spend. We ask what an all-fail task actually certifies. For the 125 tasks with no honest pass, we combine task artifacts, reference-solution runs, empty-solution controls, adversarial trials, trajectories, telemetry, and review records, and apply an ordered validity screen. Only 78 of the 125 tasks survive as certified-unsolved candidates. The remaining tasks include 14 with broken oracles, 8 dominated by infrastructure failures, 4 that are only passable through verifier bypasses, and 21 whose solvability is not certified by the available evidence. Thus, lack of saturation and genuine difficulty are not the same thing. The certified-unsolved label is also narrow: it means that the authored route passed, infrastructure did not dominate, no strict bypass was observed, and all evaluated agents failed. It does not prove intrinsic hardness, verifier completeness, or failure at the intended capability. We further analyze rejected submissions and passing tasks to show that pass rate alone cannot explain why a task is difficult. Overall, our results suggest that frontier benchmarks should report the evidence behind their all-fail tasks before using them as capability claims.