发表机构
New York University; University of California, Davis(纽约大学; 加利福尼亚大学戴维斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于大语言模型的多智能体系统贡献归因问题,提出语义合作博弈框架,定义语义夏普利值,介绍单轨迹算法SLIC,在特定基准测试中大幅降低计算成本,在多角色工作流程中表现良好,为复杂系统提供快速、无反事实且可解释的归因方法。
AI 中文摘要
贡献归因已成为基于大语言模型的多智能体系统的核心问题,现有方法存在依赖反事实评估、需重复调用模型、引入高方差且未明确捕捉中间语义状态等问题。本文提出语义合作博弈(SCG)框架,将语言流表示为语义生成超图并诱导智能体级语义价值函数,定义语义夏普利值(SSV),介绍单轨迹算法SLIC。在满足特定条件的医学基准测试中,SLIC将计算成本降低93.3%,在更一般的多角色工作流程中,SSV与扰动引起的分数下降曲线一致,揭示语义贡献和失败影响不同的情况。总体而言,SLIC为基于大语言模型的复杂多智能体系统提供了一种快速、无反事实且可解释的归因方法。
英文摘要
Contribution attribution has become a central problem in LLM-based multi-agent systems, where final outputs are produced through multiple agents, message exchanges, and ordered workflow dependencies. Existing attribution methods often rely on counterfactual valuation, such as removing agents or comparing score changes across altered agent subsets. In language-mediated workflows, these methods require repeated model calls, introduce high variance, and do not explicitly capture the intermediate semantic states through which agents produce, preserve, and transform task-relevant information. We propose Semantic Cooperative Games (SCG), a framework that represents a realized language flow as a semantic generation hypergraph and induces an agent-level semantic value function on this structure. We define the Semantic Shapley Value (SSV) to allocate contribution over semantic support logic, and introduce SLIC, a single-trajectory algorithm that constructs the semantic hypergraph, recovers minimal semantic supports, applies Boolean absorption, and computes SSV without rerunning agent subsets. We prove that SSV reduces to the classical Shapley value under standard set-based, fully observable, and no-order-dependence conditions. On a medical benchmark satisfying these conditions, SLIC reduces computation cost by 93.3% while remaining highly consistent with a Monte Carlo Shapley baseline. In more general multi-role workflows, SSV aligns with perturbation-induced score-drop profiles and exposes cases where semantic contribution and failure impact diverge. Overall, SLIC provides a fast, counterfactual-free, and interpretable attribution method for complex LLM-based multi-agent systems.