发表机构
University of Illinois at Chicago; New York University; AWS(伊利诺伊大学芝加哥分校; 纽约大学; 亚马逊云科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对小型LLM智能体,通过测试框架干预与预算感知强化学习,提升其时间预算遵守度,但发现智能体仍存在高效利用额外时间的核心挑战。
AI 中文摘要
我们研究小型大语言模型(LLM)智能体能否在明确的 wall-clock 时间预算下有效运行,既遵守分配的运行时间,又能高效利用可用时间。我们在 MLE-Bench Lite 的五个竞赛任务上评估 Qwen3.6-27B,在 Zork I(Jericho)上评估 Qwen3-4B,这两个都是智能体基准,额外的计算时间可显著提升性能。在最简单的设置中,预算仅在提示中说明,智能体无法将所述预算转化为对时间的可控使用。这些失败源于时间感知的不足,因为测试框架(harness)未提供时间反馈,同时智能体无法可靠预测动作的持续时间,也未建立从可用时间到合适策略的学习映射。我们研究两类互补的干预措施:基于测试框架的机制,用于暴露时间信息并强制截止期限;以及采用预算感知奖励的强化学习。通过测试框架注入时间信息可大幅提升 Qwen3.6-27B 的预算遵守度,且无可测量的性能损失,而强制钩子(enforcement hooks)可进一步收紧遵守度。采用 GRPO 的强化学习在 Zork I 上实现近乎完美的预算遵守度,并能泛化到训练期间未见过的保留预算,但在 MLE-Bench 上未比未训练的测试框架提升任务性能。一旦智能体被要求遵守预算,它们仍无法利用额外时间提升任务性能。经强化学习训练的策略学会了何时停止,但常以重复动作填充额外时间,而在多个预算上进行的 GRPO 训练往往会向最短预算的策略坍缩。我们的结果揭示了时间遵守与高效时间分配之间的差距,这仍是预算条件型智能体的核心挑战。
英文摘要
We study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively. We evaluate Qwen3.6-27B on five competitions from MLE-Bench Lite and Qwen3-4B on Zork I (Jericho), two agentic benchmarks where additional computational time can meaningfully improve performance. In the simplest setting, where the budget is stated only in the prompt, agents fail to translate the stated budget into controlled use of time. These failures arise from gaps in time awareness, since the harness provides no timing feedback, but also because they cannot reliably anticipate the duration of actions, and do not have a learned mapping from available time to an appropriate strategy. We investigate two complementary classes of interventions: harness-based mechanisms that expose timing information and enforce deadlines, and reinforcement learning with budget-aware rewards. Injecting timing information through the harness substantially improves budget adherence for Qwen3.6-27B without measurable loss in performance, while enforcement hooks tighten adherence further. RL with GRPO achieves near-perfect budget adherence on Zork I and generalizes to held-out budgets not seen during training, but does not improve task performance over the untrained harness on MLE-Bench. Once agents are made to respect the budget, they still fail to use additional time to improve task performance. RL-trained policies learn when to stop but often fill extra time with repeated actions, and GRPO training on multiple budgets tends to collapse toward the strategy learned for the shortest budget. Our results reveal a gap between time adherence and productive time allocation, which remains a central challenge for budget-conditioned agents.
CommentsAccepted at the NeurIPS 2026 Workshop on Resource-Aware Agentic AI. 23 pages