arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07532cs.AIcs.LGecon.TH

基于技能的智能体人工智能系统中的动态联盟形成与通信定价

Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems

Mojtaba Eslami

首次发表
浏览论文内容

中文总结 AI 辅助

针对技能型智能体 AI 系统通信低效问题,提出基于夏普利值的边际值激活与贪心路由方法,在合成实验中其效用达蛮力最优的 99.5%,激活智能体数远少于全广播。

中文摘要 AI 辅助

现代智能体人工智能系统将多个具有异构技能的大语言模型(LLM)智能体结合起来,但大多数架构要么预先固定通信方式,要么允许全广播。这两种方式都可能效率低下,因为 token 成本、延迟、冗余度和错误传播会随着活跃智能体和通信链路数量的增加而上升。我们将智能体选择与通信建模为具有任务条件净效用 $U(C\mid x)=V(C\mid x)-\sum_{i\in C}c_i$ 的合作博弈,将联盟层面成本与智能体激活成本分离。我们提出边际值激活规则和贪心路由器,扩展模型以优化具有每条边成本的通信边,并使用估计的夏普利值(Shapley values)在执行前和执行期间预测哪些智能体值得联系。我们将该问题与子模最大化关联起来,并证明两个有限保证:针对单调、基数约束特殊情况的曲率优化边界,以及针对无约束非单调情况通过双贪心(double greedy)得到的、带符号目标修正的紧 1/2 近似。这两个保证都不直接适用于主路由器,主路由器仍为启发式方法。我们还证明了夏普利-子模夹层边界,将边际值路由的误差与每个智能体的边际收益递减量关联起来。在合成实验中,贪心路由器达到了蛮力最优效用的 99.5%,同时平均激活 8 个智能体中的 1.96 个,相比之下全广播仅为 38.8%。性能对激活成本和冗余权重具有鲁棒性,但在子模性严重违反或价值估计有噪声时降至 66%。我们将该框架与夏普利定价、享乐联盟形成和通信图剪枝区分开来,并提出在真实多智能体 LLM 基准上进行评估。

英文摘要

Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in advance or allow full broadcast. Both can be inefficient because token cost, latency, redundancy, and error propagation increase with the number of active agents and communication links. We model agent selection and communication as a cooperative game with task-conditioned net utility $U(C\mid x)=V(C\mid x)-\sum_{i\in C}c_i$, separating coalition-level costs from agent activation costs. We propose a marginal-value activation rule and greedy router, extend the model to optimize communication edges with per-edge costs, and use estimated Shapley values to predict which agents are worth contacting before and during execution. We connect the problem to submodular maximization and prove two limited guarantees: a curvature-refined bound for a monotone, cardinality-constrained special case, and a tight $1/2$-approximation, with a correction for signed objectives, for an unconstrained non-monotone case via double greedy. Neither guarantee applies directly to the main router, which remains a heuristic. We also prove a Shapley-submodularity sandwich bound linking the error of marginal-value routing to a per-agent diminishing-returns quantity. In synthetic experiments, greedy routing achieves $99.5%$ of brute-force-optimal utility while activating $1.96$ of $8$ agents on average, compared with $38.8%$ for full broadcast. Performance is robust to activation cost and redundancy weight but falls to $66%$ under strong violations of submodularity or noisy value estimates. We distinguish the framework from Shapley pricing, hedonic coalition formation, and communication-graph pruning, and propose evaluation on real multi-agent LLM benchmarks.

发表机构

  • University of Calgary(卡尔加里大学)

机构由 AI 辅助整理,请以论文原文为准。

↑