发表机构
The University of Hong Kong; Hunyuan Team, Tencent(香港大学; 腾讯混元团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
$R^3$-Bench测试发现,大语言模型在共享预算下的资源理性推理存在不足,其表现出的单问题能力与共享预算下的实际表现存在差距。
AI 中文摘要
在认知科学中,资源理性研究的是智能体应如何分配有限的计算资源以最大化期望价值。大多数推理和智能体基准测试采用每个任务独立的预算;现有的共享预算研究未将基准测试套件的性能与同一模型在单个问题上表现出的能力进行校准。我们推出$R^3$-Bench,该基准测试套件在无工具和智能体两种场景下,针对数学、竞赛编程和抽象推理三个领域,评估共享预算下的六问题套件。匹配的单问题响应曲线定义了基于观测成功的离线经验神谕。在六个模型的72个主表单元格中,神谕均值在所有单元格中均达到或超过竞赛均值,且在71个单元格中严格更高。在中等无工具压力下,均等分配重放也使六个模型中的四个超过竞赛性能。轨迹诊断显示策略更新有限及压力依赖的失败模式。在强智能体压力下的三个模型诊断中,至少一个固定调度器在九个单元格中的六个超过竞赛均值,但没有策略在所有领域占主导。这些结果揭示了已展现能力与共享预算实现之间的持续差距。
英文摘要
In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same model's demonstrated single-problem competence. We introduce $R^3$-Bench, which evaluates six-problem suites under shared budgets across mathematics, competitive programming, and abstract reasoning in tool-free and agentic settings. Matched single-problem response curves define an offline empirical oracle over observed successes. Across 72 main-table cells for six models, the oracle mean matches or exceeds the contest mean in all cells and is strictly higher in 71. Under moderate tool-free pressure, equal-allocation replay also exceeds contest performance for four of six models. Trajectory diagnostics reveal limited strategy updating and pressure-dependent failure patterns. In a three-model diagnostic under strong agentic pressure, at least one fixed scheduler exceeds the contest mean in six of nine cells, but no policy dominates across domains. These results expose a persistent gap between demonstrated competence and shared-budget realization.
CommentsCode is available at https://github.com/NineAbyss/R-3-Bench . The dataset is available at https://huggingface.co/datasets/R-3-Bench/R-3-Bench