发表机构
Shanghai Jiao Tong University; Eastern Institute of Technology; University of Science and Technology of China; Ningbo Institute of Digital Twin(上海交通大学; 宁波诺丁汉大学; 中国科学技术大学; 数字孪生宁波研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对在线智能体自我进化与当前服务竞争GPU资源的问题,提出状态感知调度器LearnSched,通过统一投资模型和一步反事实回放选择进度、检查点或等待,实验表明可提升净价值并简化调度。
AI 中文摘要
在线智能体可以通过构建可复用的工具、指南或模型状态来改进未来的服务,但这些工作与当前请求竞争同一批GPU资源。利用空闲计算进行自我进化面临一个根本性的系统约束:收益仅在工件发布并被使用之后才能实现,而暂停进化则会让服务容量等待内存释放和运行时恢复。一项值得完成的投资因此可能值得推迟。我们提出了LearnSched,一种状态感知的调度器,它将能力(capability)的重用窗口和计算的及时返回纳入一个统一的投资模型。LearnSched将候选进度、检查点开销和当前恢复路径纳入动作值。在相同的信息和容量约束下,它使用一步反事实回放(one-step counterfactual rollout)来选择进度推进、检查点保存或等待,相对于一个完全计费窗口策略。我们通过独立的A100组件测量来表征交接成本,并在1920个配对有限模型场景中评估动作选择。当冷恢复成本高昂时,等待更长的执行窗口可以通过避免交接来提高净价值;在评估的热恢复场景中,强窗口策略已经实现了相同的价值。保留可恢复状态可以缩短容量返回时间,并减少对复杂进化调度的需求。
英文摘要
Online agents can improve future service by constructing reusable tools, guidance, or model states, but this work competes with current requests for the same GPUs. Exploiting idle compute for self-evolution faces a fundamental systems constraint: benefits arrive only after an artifact is published and used, while pausing evolution leaves service capacity waiting for memory release and runtime recovery. An investment worth completing may therefore be worth postponing. We present LearnSched, a state-aware scheduler that brings the reuse window of a capability and the timely return of compute into a common investment model. LearnSched incorporates candidate progress, checkpoint overhead, and the current recovery path into action values. Under the same information and capacity constraints, it uses one-step counterfactual rollout to choose progress, checkpointing, or waiting relative to a fully costed window policy. We characterize handoff costs through independent A100 component measurements and evaluate action selection in 1920 paired finite-model scenarios. When cold recovery is expensive, waiting for longer execution windows can improve net value by avoiding handoffs; in evaluated warm-recovery scenarios, the strong window policy already achieves the same value. Retaining recoverable state can shorten capacity return and reduce the need for complex evolution scheduling.
Comments21 pages, 17 figures