发表机构
Andromede AI(安卓迈德人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对17个跨多领域的智能体基准,提出无需展开的任务难度预测方法,发现token级熵是有效预测信号,可通过难度残差暴露环境缺陷,助力校准基准与构建训练课程。
AI 中文摘要
任务难度决定智能体的成功概率,无需展开(rollouts)进行估算意味着在有状态环境中执行高成本模拟前,直接从任务描述预测难度。可靠的估算因此能让环境设计者校准评估基准并构建渐进式训练课程,随着智能体进入长视野领域,经验试错成为严重的计算瓶颈,这一点愈发重要。早期预测的现有工作局限于静态任务或孤立编码环境,常依赖狭窄特征和不准确评估指标。我们研究了涵盖编码、数学、机器学习、网页导航、函数调用等领域的17个智能体基准的事前(ex ante)难度预测。我们表明AUC可能掩盖糟糕的难度估算,确定 token 级熵是有用的预测信号,并展示预期与观测难度间的残差如何暴露隐藏的环境缺陷,如污染和不可行性。
英文摘要
A fundamental challenge in evaluating and training autonomous agents is measuring the intrinsic difficulty of the tasks they attempt. Estimating this quantity can be useful for environment designers creating synthetic data or benchmarks, as well as for constructing training curricula, thereby offering a way to reduce compute costs. Such an estimation becomes increasingly important as agents (particularly LLM-based) move into longer-horizon domains, where empirical trial-and-error becomes a severe computational bottleneck. However existing work is largely confined to static tasks and relies on evaluation metrics that, as we show, can give an incomplete picture of predictive performance. In this paper we study ex ante difficulty prediction across 17 agentic benchmarks spanning coding, mathematics, machine learning, web navigation, function calling, and other domains. We find that AUC can mask poor difficulty estimates, identify token-level entropy as a useful predictive signal, and demonstrate how residuals between expected and observed difficulty can expose environment flaws such as contamination and infeasibility.