AI 中文总结
该研究针对推理大语言模型的测试时缩放,明确三种推理机制,提出系统评估原则与可复现性要求,应用于多类基准并发布超20亿条推理轨迹。
AI 中文摘要
大语言模型可利用更多推理时计算求解难度显著更高的推理问题。然而,“测试时缩放”这一术语如今涵盖多种推理算法,包括沿单一轨迹扩展思考过程、采样完整候选并通过投票或验证聚合结果、对未完成的部分状态进行搜索等。这些算法在统计结构、计算核算及失败模式上存在差异。若将这些程序在单一标量“预算”下视为可互换,或报告未说明推理协议的准确率,会导致不同研究间的结果难以比较。我们沿三个维度构建了测试时缩放的系统说明:首先,将测试时缩放形式化为自回归模型隐式前缀树上的预算推理,并区分三种结构机制:单轨迹顺序缩放、带终端归约的叶级缩放、前缀级缩放;其次,将评估对象视为整个推理系统,开发了将端到端系统性能与候选库诊断分离的评估原则,引入一种评估轮廓,其坐标和简单泛函可恢复或界定常见的重复采样指标,并规定与协议匹配的计算和不确定性报告方式;第三,明确推理协议的可复现性要求,区分精确重放与分布可复现性,并确定支持每种要求所需的人工制品。我们还按模型侧和接口机制梳理了开放权重推理生态系统,将这些原则应用于广泛知识、符号推理和竞赛数学基准,并组装了超过20亿条完整推理轨迹,以提供逐步丰富的验证器和 token 级信号。
英文摘要
Large language models can solve harder reasoning problems with more inference-time compute. The term "test-time scaling," however, covers several inference algorithms: extending deliberation along one trajectory, sampling completed candidates and aggregating them by voting or verification, and searching over partial states. These algorithms differ in statistical structure, compute requirements, and failure modes. Treating them as interchangeable under a scalar "budget," or reporting accuracy without specifying the inference protocol, makes results difficult to compare across studies. We study test-time scaling along three axes. First, we formalize it as budgeted inference over the implicit prefix tree of an autoregressive model and distinguish single-trajectory sequential scaling, leaf-level scaling with terminal reduction, and prefix-level scaling. Second, we treat the full inference system as the evaluated object and separate end-to-end performance from candidate-bank diagnostics. We introduce an evaluation profile whose coordinates and simple functionals recover or bound common repeated-sampling metrics, and require compute accounting and uncertainty estimates that match the protocol. Third, we distinguish exact replay from distributional reproducibility and state the requirements for each. We also organize open-weight reasoning models by model-side and interface mechanisms. Our empirical study covers broad knowledge, symbolic reasoning, and competition mathematics, and we publicly release 1,403,520 sampled model attempts. The project website is available at https://mohsenhariri.github.io/scorio/tts. The released datasets are Trace (https://huggingface.co/datasets/harimo/scorio-trace), Lite (https://huggingface.co/datasets/harimo/scorio-lite), Math (https://huggingface.co/buckets/harimo/scorio-math), and SuperGPQA (https://huggingface.co/buckets/harimo/scorio-gpqa).