arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13890cs.MAcs.SE

学习协作程度:面向多智能体代码生成的难度感知拓扑选择

Learning How Much to Collaborate: Difficulty-Aware Topology Selection for Multi-Agent Code Generation

Yunsong Hong

首次发表
浏览论文内容

中文总结 AI 辅助

针对多智能体代码生成中固定拓扑粒度不当的问题,提出难度感知拓扑选择器(DATS),通过图网络预测各拓扑成功率并平衡成本,在预算匹配下显著提升性能。

中文摘要 AI 辅助

多智能体代码生成系统部署时使用单一通信拓扑,且针对每个问题仅选择一次。这种粒度是错误的。我们在来自APPS、HumanEval+和LiveCodeBench的614个问题上评估了五种拓扑,发现层级协作相对于单智能体的优势从最容易三分之一问题上的2.4个pass@1百分点增长到最难三分之一问题上的21.1个百分点,而其令牌成本始终高出约十倍。我们提出了难度感知拓扑选择器(DATS),它预测每种拓扑解决某个问题的概率,并选择最大化预测成功减去成本的那个。其预测器是一个图网络,将五种拓扑视为连通性顺序的节点而非独立标签,相比扁平多标签头提升了1.7个百分点。由于成本惩罚是一个无需重新训练即可重新校准的单一标量,路由器在同等花费下进行比较:在这种预算匹配协议下,六种成本感知方法跨越了21.6个百分点,而两个领先于DATS的基线在根据该协议校准后落后。固定为始终层级成本的40%时,DATS达到77.7%的pass@1,而始终层级为73.6%,最强学习型竞争者为74.3%,所有十一个成对McNemar比较均通过Holm-Bonferroni校正。这4.1个百分点的提升在跨越十四个能力百分点的四个骨干网络上保持一致,将39个可解释特征替换为图网络或预训练编码器最多使准确率变化1.3个百分点,且从未显著。一项针对400个数学推理问题的跨领域研究重现了这一效应,差距从2.5个百分点扩大到20.9个百分点。

英文摘要

Multi-agent systems for code generation are deployed with a single communication topology, chosen once for every problem. This is the wrong granularity. Evaluating five topologies on 614 problems from APPS, HumanEval+ and LiveCodeBench, we find that the advantage of hierarchical collaboration over a single agent grows from 2.4 points of pass@1 on the easiest third of problems to 21.1 points on the hardest third, while its token cost stays about ten times higher. We propose the Difficulty-Aware Topology Selector (DATS), which predicts each topology's probability of solving a problem and selects the one maximising predicted success minus cost. Its predictor is a graph network that treats the five topologies as nodes of a connectivity order rather than independent labels, worth 1.7 points over a flat multi-label head. Because the cost penalty is a single scalar recalibrable without retraining, routers compare at equal spend: under this budget-matched protocol six cost-aware methods span 21.6 percentage points, and two baselines leading DATS fall behind once calibrated to it. Fixed at 40% of the always-hierarchical cost, DATS reaches 77.7% pass@1 against 73.6% (always-hierarchical) and 74.3% (strongest learned competitor), all eleven pairwise McNemar comparisons surviving Holm-Bonferroni correction. The 4.1-point gain holds across four backbones spanning fourteen points of capability, and replacing the 39 interpretable features with a graph network or a pretrained encoder shifts accuracy by at most 1.3 points, never significantly. A cross-domain study on 400 mathematical reasoning problems reproduces the effect, the gap widening from 2.5 to 20.9 points.

发表机构

  • School of Computer Science, The University of Sydney(悉尼大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑