arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07257cs.AI

MemMux:并行编码智能体集群的运行时验证与诚实资源归属

MemMux: Runtime Verification and Honest Resource Attribution for Fleets of Parallel Coding Agents

Sumanyu Muku

首次发表
浏览论文内容

中文总结 AI 辅助

MemMux 通过运行时验证将资源治理转化为可检查信号,实现并行编码智能体集群的精确内存归属、完全回收与逃逸进程可见性,在受限预算下零交换运行。

中文摘要 AI 辅助

开发者越来越多地在一台工作站上并行运行一组编码智能体。他们使用的工具,如终端复用器 tmux 和新一代智能体管理器,是为排列窗口而构建的,而非管理内存。当十个智能体各自生成语言服务器、测试运行器和浏览器时,没有标准工具能说明每个智能体占用多少内存,确认被终止的智能体的后代进程已完全消失,发现逃逸出其智能体的子进程,或在内存不足(OOM)杀死进程会静默丢弃未提交工作时使机器远离交换空间悬崖。我们将这些问题视为运行时验证问题:智能体托管基础平台应在智能体运行时持续发出操作员或审计员可检查的可观察信号。我们提出 MemMux,一个本地运行时,将资源治理转化为可检查的信号(每个智能体的资源归属、完全回收、逃逸进程可见性、超量使用下的有界占用及监控开销),并附带一个基于声明约束的基准测试,与 tmux、一个专用智能体复用器及原始进程基线在相同工作负载下进行对比。在 Linux 主机上,在受限内存预算下,MemMux 通过准入子集并在压力下回收资源,使集群保持在预算(7.5 GiB)内且零交换,而未治理的工具运行所有智能体,将机器钉在内存上限(超出预算 2 倍),并溢出约 2 GiB 到交换空间。MemMux 能回收被终止智能体进程子树的 100%,而原始基线会遗留一半,且只有它能发现逃逸子进程(10 个中检测到 10 个)。我们报告了成本:1 Hz 归属扫描在单个智能体时占用约 0.6% CPU,但在十个智能体时占用 2.7%,高于我们 2% 的目标。在真实 Claude Code 会话上运行测试工具显示,100% 归属和低开销可延续到实时智能体树。我们发布了引擎、基准测试及一条命令即可复现的工具。

英文摘要

Developers increasingly run a fleet of coding agents side by side on one workstation. The tools they reach for, terminal multiplexers like tmux and a new generation of agent managers, were built to arrange windows, not to govern memory. When ten agents each spawn language servers, test runners, and browsers, no standard tool can say how much memory belongs to which agent, confirm that a terminated agent's descendants are gone, notice a child that has escaped its agent, or keep the machine off the swap cliff when an OOM kill would silently discard uncommitted work. We treat these as runtime-verification problems: an agent-hosting substrate should continuously emit observable signals an operator or auditor can check while agents run. We present MemMux, a local runtime that turns resource governance into checkable signals (per-agent attribution, complete reclamation, escaped-process visibility, bounded footprint under overcommit, and monitoring overhead), with a claims-disciplined benchmark against tmux, a purpose-built agent multiplexer, and a raw-process baseline on identical workloads. Under a binding memory budget on a Linux host, MemMux keeps the fleet under budget (7.5 GiB) with zero swap by admitting a subset and reclaiming under pressure, while the ungoverned tools run every agent, pin the machine at its RAM ceiling (2x over budget), and spill about 2 GiB into swap. MemMux reclaims 100% of a terminated agent's process subtree where the raw baseline strands half of it, and it alone surfaces escaped children (10 of 10 detected). We report the cost: the 1 Hz attribution scan runs near 0.6% CPU at one agent but 2.7% at ten, above our 2% target. Running the harness on real Claude Code sessions shows 100% attribution and low overhead carry over to live agent trees. We release the engine, benchmark, and a one-command reproducer.

发表机构

  • Amira Learning

机构由 AI 辅助整理,请以论文原文为准。

↑