发表机构
Carnegie Mellon University; AWS AI Labs; Washington University in St. Louis; Georgia Institute of Technology(卡内基梅隆大学; AWS AI 实验室; 圣路易斯华盛顿大学; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Hermes框架和Hermes-Learn训练方法,将上下文分配决策转移给模型,使模型能利用额外推理计算进行扩展,并提升跨基准的泛化能力。
AI 中文摘要
测试时扩展通过推理过程中分配额外计算来提升模型性能。要在多个上下文窗口之间有效利用这些计算,需要决定如何分配新的上下文以及在其间携带哪些信息。我们将模型做出这些决策的能力称为上下文推理。现有方法主要通过其框架来预设这些决策;我们则将这些决策转移给模型本身。我们引入了1)Hermes,一系列简单、可配置的框架,逐步变化模型对上下文分配和重用的控制;以及2)Hermes-Learn,一个两阶段框架,用于学习这些能力。我们发现,能力强的模型可以利用这种灵活性,随额外推理时间计算进行扩展,而较小的开源模型最初难以做到这一点。使用Hermes-Learn进行训练弥补了这一差距,诱导出随问题和推理进展而变化的适应性上下文推理策略。这些收益在多个基准和模型中普遍存在,外推到训练期间未见过的推理时间计算,并迁移到Hermes之外的互补测试时扩展方法。
英文摘要
Test-time scaling improves model performance by allocating additional compute during inference. Using this compute effectively across multiple context windows requires deciding how to allocate fresh contexts and what information to carry between them. We call a model's ability to make these decisions contextual reasoning. Existing approaches largely prescribe these decisions through their harness; we instead shift them to the model. We introduce 1) Hermes, a family of simple, configurable harnesses that progressively varies model control over context allocation and reuse, and 2) Hermes-Learn, a two-stage framework for learning these capabilities. We find that capable models can exploit this flexibility to scale with additional inference-time compute, while smaller open-source models initially struggle to do so. Training with Hermes-Learn closes this gap, inducing adaptive contextual reasoning strategies that vary with both the problem and the progress of reasoning. These gains generalize across benchmarks and models, extrapolate beyond the inference-time compute seen during training, and transfer to complementary test-time scaling methods beyond Hermes.
Comments45 pages