arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02750cs.AI

双层协调反思:多智能体大语言模型系统的博弈论方法

Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems

Yihang Chen, Yuxiang Chen, Yuxuan Huang, Meng Fang, Weilin Luo, Jun Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多智能体LLM系统缺乏统一协调解释的问题,提出博弈论模型并引入SRMA算法,实验显示基于Kimi的系统在SWE-bench上解决率达72.2%,优于基准系统。

中文摘要 AI 辅助

多智能体大语言模型(LLM)系统通常采用一个协调器为一组工作智能体分解任务,随后通过文本反思进行改进。尽管这些系统取得了良好的实验结果,但它们缺乏对协调、记忆改进以及外部验证作用的统一解释。我们将协调器与工作智能体的交互建模为一个双层协调博弈:在有限耦合下,工作智能体的局部更新博弈是一个近似势博弈,其均衡松弛由分解质量控制。接着,我们将反思分析为语义记忆状态上的随机移动。对于自由形式的反思,我们推导了有限时间上界,证明了最坏情况的紧性,并在可证伪的持续谐波条件下给出了正下界。我们进一步证明了一个信息论不可能性结果:仅观察生成的文本记录的门无法在文本不可区分的环境中实现均匀改进,而基于环境的门则可以。受此分离的启发,我们引入了随机反思记忆上升(SRMA),它仅在基于环境的评估风险严格降低时才接受候选记忆。在校准和非退化校正质量下,SRMA可精确收敛、几何收敛或多项式收敛;匹配构造表明,两种收敛速率 regime 均为阶紧。我们还为随机评估提供了置信门控,并为分段平稳环境提供了重锚定保证。实验使用基于环境的指标实例化这些对象,并测试了预测的协调和漂移规律。在500个SWE-bench实例上,完整的基于Kimi的系统解决了72.2%的问题,而公开的mini-SWE-agent基准的解决率为70.8%。代码:this https URL

英文摘要

Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memory improvement, and the role of external verification. We model orchestrator-worker interaction as a bilevel coordination game: under bounded coupling, the workers' local-update game is an approximate potential game whose equilibrium slack is controlled by decomposition quality. We then analyse reflection as stochastic movement over semantic memory states. For free-form reflection, we derive a finite-time upper bound, prove worst-case tightness, and give a positive lower bound under a falsifiable persistent-harm condition. We further prove an information-theoretic impossibility result: no gate that observes only the generated transcript can improve uniformly over text-indistinguishable environments, whereas an environment-grounded gate can. Motivated by this separation, we introduce Stochastic Reflective Memory Ascent (SRMA), which accepts a candidate memory only after a grounded evaluation risk strictly decreases. Under calibration and non-degenerate corrective mass, SRMA converges exactly, geometrically or polynomially; matching constructions show that both rate regimes are order-tight. We also provide confidence gating for stochastic evaluation and re-anchoring guarantees for piecewise-stationary environments. Experiments instantiate these objects with environment-grounded metrics and test the predicted coordination and drift laws. On 500 SWE-bench instances, the complete Kimi-based system resolves 72.2% versus a 70.8% public mini-SWE-agent reference. Code: https://github.com/YihangChen9/Bilevel-Coordinated-Reflection

发表机构

  • UCL Centre for Artificial Intelligence(伦敦大学学院人工智能中心)
  • University of Liverpool(利物浦大学)
  • Huawei(华为)

机构由 AI 辅助整理,请以论文原文为准。

↑