发表机构
Meituan LongCat Team; Zhejiang University; Westlake University; Allen Institute for AI; University of Washington(美团龙猫团队; 浙江大学; 西湖大学; 艾伦人工智能研究所; 华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过分析AutoLab和EdgeBench基准,提出能力-时间扩展模型,并推导出继续策略,在节省约三分之一计算时间的同时,将得分损失控制在2.4%-3.3%以内。
AI 中文摘要
长时程智能体通过持续交互、执行和任务反馈来改进解决方案。扩展研究将性能与资源和能力相关联,然而现有能力如何影响长时间交互的回报仍未被充分理解。为弥补这一空白,我们分析了AutoLab和EdgeBench这两个长时程基准。我们发现,起始性能和后续增长与不同的能力相关联:在任务类别内,相似的早期得分可能预示着不同的后期增益。为形式化这一发现,我们建立了能力-时间扩展模型,采用跨模型共享的类别特定逻辑幂律。拟合到早期轨迹后,这些曲线将观察到的模型类别平均得分外推到后续计算阶段。然而,平均得分的上升掩盖了改进机会的收窄:后期增益集中在少数改进中的模型上。高最终得分和持续改进也具有不同的能力特征。预测的平均增益估计每个模型改进任务的占比;对这些估计取平均可预测改进模型的平均占比。这些不均衡的回报促使我们决定是否应继续某个特定运行。因此,我们推导出一个继续策略,以节省时间和计算资源,同时将得分损失限制在较小范围内。该策略根据运行中观察到的进度对增长预测进行条件化,并权衡即时与延迟增益及计算成本。在基于其他模型历史进行训练和价格校准的回放中,该策略节省了约三分之一的完整运行时间,相对得分损失在AutoLab单次运行上为2.4%,在EdgeBench公布的均值曲线上为3.3%。我们的代码库可在https://this URL获取。
英文摘要
Long-horizon agents improve solutions through sustained interaction, execution, and task feedback. Scaling studies relate performance to resources and capabilities, yet how existing capabilities shape returns to extended interaction remains less understood. To address this gap, we analyze AutoLab and EdgeBench, two long-horizon benchmarks. We find that starting performance and subsequent growth are associated with different capabilities: within a task category, similar early scores can precede different later gains. To formalize this finding, we model capability-time scaling with category-specific logistic power laws shared across models. Fitted to early trajectories, these curves extrapolate the observed models' category-average scores to later computation. However, rising average scores mask narrowing improvement opportunities: later gains concentrate among fewer improving models. High final scores and continued improvement also have distinct capability profiles. Predicted mean gains estimate each model's fraction of improving tasks; averaging these estimates forecasts the average share of improving models. These uneven returns motivate deciding whether a specific run should continue. We therefore derive a continuation policy to save time and compute with limited score loss. The policy conditions growth predictions on the run's observed progress and weighs immediate and delayed gains against computation costs. In replay with training and price calibration based on other models' histories, the policy saves roughly one-third of full-run time, with relative score losses of 2.4% on AutoLab individual runs and 3.3% on EdgeBench published mean curves. Our repository is available at https://github.com/Chihaya-Anon-chan/long-horizon-scaling.