arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高扇出智能体沙箱的内存压缩

Memory Compression for High-Fanout Agent Sandboxes

Mengming Li, Ceyu XU, Qijun Zhang, Jiangnan Yu, Xiangfeng Sun, Haohui Mai, Zhiyao Xie

arXiv 2609.11294首次发表:更新:

发表机构

HKUST(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对高扇出智能体沙箱的内存冗余问题,提出AgentZip压缩系统,利用模板与跨沙箱冗余、扩大压缩范围并配合恢复预取和LLM等待调度,将内存减少达8.7倍,减速降至1.40倍。

AI 中文摘要

高扇出智能体工作负载会形成日益严重的内存瓶颈,因为单个任务可能产生大量并发沙箱会话。然而,这些沙箱远非相互独立:它们源自共享模板并执行相关轨迹,暴露出大量的模板相对冗余和跨沙箱内存冗余。传统内存压缩在三个基本维度上与这种场景不匹配:如何压缩,因为它们未能利用非相同沙箱页面之间的相似性;压缩什么,因为它们通过保守的页面选择来控制缺页开销;以及何时压缩,因为压缩要么由内存压力触发,要么在不了解智能体执行阶段的情况下进行。我们提出了AgentZip,这是首个专为AI智能体沙箱设计的内存压缩系统。AgentZip引入了利用模板相对冗余和跨沙箱冗余的压缩机制。它将压缩范围扩大到任何具有有利表示的页面,并将开销控制从压缩时的页面选择转移到恢复时的预取。它进一步将昂贵的压缩与LLM等待时段对齐,以避免干扰前台工具执行。在LLM训练和推理工作负载中,AgentZip将沙箱拥有的内存减少了高达8.7倍,而Linux配置仅为2.1倍。恢复预取和智能体执行感知调度将激进压缩的减速从高达3.1倍降低到1.40倍,同时保留了几乎所有的内存节省收益。

英文摘要

High-fanout agent workloads create a growing memory bottleneck because a single task may spawn many concurrent sandbox sessions. Yet these sandboxes are far from independent: they originate from a shared template and execute related trajectories, exposing substantial template-relative and cross-sandbox memory redundancy. Conventional memory compression is poorly matched to this setting in three fundamental dimensions: how to compress, because they fail to exploit similarity across non-identical sandbox pages; what to compress, because they control page-fault overhead through conservative page selection; and when to compress, because compression is either triggered by memory pressure or performed without awareness of agent execution phases. We present AgentZip, the first memory compression system designed specifically for AI-agent sandboxes. AgentZip introduces compression mechanisms that exploit both the template-relative and cross-sandbox redundancy. It broadens the compression scope to any page with a profitable representation and shifts overhead control from compression-time page selection to restore-time prefetching. It further aligns expensive compression with LLM waiting periods to avoid interfering with foreground tool execution. Across LLM training and inference workloads, AgentZip reduces sandbox-owned memory by up to 8.7x, compared with 2.1x for the Linux configuration. Restore prefetching and agent-execution-aware scheduling reduce the slowdown of aggressive compression from as high as 3.1x to 1.40x while retaining nearly all of its memory-saving benefit.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑