发表机构
Beihang University; University of Leeds; The University of Sydney(北京航空航天大学; 利兹大学; 悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型智能体服务中沙盒部署的资源利用与延迟矛盾,提出SpecBox,通过推测性沙盒预分配、上下文感知随机预取及相关优化,有效降低端到端延迟并削减峰值内存消耗。
AI 中文摘要
随着大语言模型智能体越来越依赖模型上下文协议调用隔离的外部沙盒,解耦的沙盒部署在资源利用和交互式尾部延迟之间引入了基本矛盾。持久的长期沙盒预留会产生过高的内存开销,而按需延迟实例化会导致严重的冷启动惩罚。为解决这一困境,我们提出了SpecBox,它围绕为动态大语言模型智能体执行管道量身定制的推测性沙盒预分配构建。其核心实现了关键字匹配和流语义嵌入以实现意图驱动的沙盒预热,利用上下文感知随机预取扩展预热窗口,还通过语义结果缓存和专用带外共享内存传输平面进行优化。在高并发多轮智能体跟踪上评估,原型表明SpecBox将P99端到端延迟相对于按需沙盒基线最多降低2.9倍,同时将峰值内存消耗与永久预留沙盒部署相比削减45.9%。
英文摘要
As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency. Persistent long-lived sandbox reservations incur excessive memory overhead at scale, while lazy on-demand instantiation generates severe cold-start penalties that degrade response performance under multi-tenant, multi-turn agent workloads. To resolve this dilemma, we present SpecBox, a runtime built around speculative sandbox preallocation tailored for dynamic LLM agent execution pipelines. At its core, SpecBox implements keyword matching and streaming semantic embedding to enable intent-driven sandbox prewarming, which identifies pending tool execution demands mid-LLM token generation and fully overlaps sandbox bootstrapping with model inference. To extend prewarming windows across sequential agent steps, the framework leverages context-aware stochastic prefetching atop a sandbox dependency graph to probabilistically forecast future sandbox switches ahead of execution. We complement these speculative mechanisms with two orthogonal optimizations: a semantic result cache that prunes redundant repeated sandbox invocations, and a dedicated out-of-band shared-memory transport plane that bypasses conventional network serialization to deliver zero-copy artifact transfers. Evaluated on high-concurrency multi-turn agent traces, our prototype demonstrates that SpecBox cuts P99 end-to-end latency by up to $2.9\times$ relative to the on-demand sandbox baseline, while slashing peak memory consumption by $45.9\%$ compared to permanently reserved sandbox deployments.