发表机构
Toga Networks (Huawei)(托加网络(华为))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多智能体大语言模型系统运行时无自我重组机制的问题,提出自主拓扑变异(ATM)机制,结合遥测驱动过载检测与安全不变量,经实验验证可提升任务成功率、减少隐私内存暴露并降低延迟,且相关内容已开源。
AI 中文摘要
多智能体大语言模型框架通常在启动时固定其团队拓扑。当单个智能体在运行时过载,例如混合过多动作类别、累积工具错误或在过多调用后排队时,系统没有自我重组机制。我们引入自主拓扑变异(ATM),一种用于多智能体大语言模型框架的运行时团队变异机制。ATM将遥测驱动的过载检测与三个安全不变量相结合,这些不变量控制每个结构变化:能力单调性、状态路由完整性和实时前影子验证。ATM监控一个六信号瓶颈指数,包括队列深度、上下文抖动、工具错误率、角色熵、重试循环率和跨智能体等待时间。当连续多个滴答超过预热校准阈值时,ATM将过载智能体分解为专门的子智能体,并将父智能体热交换为协调器角色,同时保留其外部身份。状态转移由隐私级别感知路由控制:每个内存原子仅路由到允许的子智能体集,或根据记录的原因明确丢弃。在通过影子验证窗口之前,没有候选拓扑接收实时流量。在四个消融条件和三个工作负载下,使用确定性工具存根进行720次由DeepSeek-V3驱动的任务运行,ATM分解提升将代码任务成功率从3.3%提高到61.7%。完整的轨道和蒸馏系统将正则表达式分类器下检测到的高隐私内存暴露从每个任务2.0次事件减少到0.0次事件,同时保持任务质量。承载ATM不变量的运行时轨道在智能体热路径上增加的p99延迟不到500微秒。包括一个带有真实Python执行的小型实时工具探针作为外部有效性检查。实现、基准测试工具和跟踪信息已开源。
英文摘要
Multi-agent LLM frameworks typically fix their team topology at boot time. When an individual agent becomes overloaded at runtime, for example by mixing too many action categories, accumulating tool errors, or queueing behind too many calls, the system has no mechanism to restructure itself. We introduce Autonomous Topology Mutation (ATM), a runtime team-mutation mechanism for multi-agent LLM frameworks. ATM combines telemetry-driven overload detection with three safety invariants that gate each structural change: capability monotonicity, state-routing completeness, and shadow-before-live validation. ATM monitors a six-signal Bottleneck Index that includes queue depth, context thrash, tool-error rate, role entropy, retry-loop rate, and cross-agent wait time. When a warmup-calibrated threshold is breached for multiple consecutive ticks, ATM factorises the overloaded agent into specialised sub-agents and hot-swaps the parent into a coordinator role while preserving its external identity. State transfer is controlled by privacy-level-aware routing: each memory atom is routed only to a permitted child set, or explicitly dropped with a logged reason. No candidate topology receives live traffic until it has passed a shadow validation window. On 720 DeepSeek-V3-driven task runs with deterministic tool stubs across four ablation conditions and three workloads, the ATM factoriser split lifts code-task success from 3.3% to 61.7%. The full rail-and-distillation system reduces detected high-privacy memory exposure under a regex classifier from 2.0 to 0.0 events per task while preserving task quality. The runtime rails carrying ATM's invariants add less than 500 microseconds of p99 latency on the agent hot path. A small live-tool probe with real Python execution is included as an external-validity check. The implementation, benchmark harness, and traces are open-sourced.
Comments9 pages, 5 tables. Code and benchmark harness available at https://github.com/sidikbro/jiuwen_atm