DART:一种基于DAG的声誉与激励机制框架,通过区块链赋能的治理实现可信的LLM多智能体协作
DART: A DAG-Based Reputation and Incentive Framework via Blockchain-Enabled Governance for Trustworthy LLM Multi-Agent Collaboration
浏览论文内容
中文总结 AI 辅助
DART提出基于DAG的声誉与激励框架,结合区块链治理,实现可信LLM多智能体协作,在GSM8K达93.6%准确率,并有效隔离恶意智能体。
中文摘要 AI 辅助
基于大型语言模型(LLM)的多智能体系统(MAS)主要依赖集中式编排,且缺乏针对智能体可靠性、参与度及系统级行为一致性的正式验证机制。这些缺陷使得开放环境极易受到不合作或恶意智能体的攻击。本文提出DART,一种基于有向无环图(DAG)的声誉与激励调节框架,用于可信的多智能体协作,将集中式操作编排与区块链赋能的去中心化治理和问责制相结合。DART统一了DAG工作流编排、能力和声誉感知的任务分配、动态行为更新、多因素激励,以及智能合约问责与IPFS存储。在该范式下,智能体选择动态平衡任务匹配度、历史声誉和工作负载,而执行后的行为证据持续校准智能体的信任度及未来参与概率。在四个维度上的评估中,DART在GSM8K上达到93.6%的Pass@1,并使用两个智能体在142秒内构建了一个全栈应用,优于集中式基线。在五次独立的150轮纵向试验中,完整版DART实现了93.33±2.26%的平均任务成功率、0.9357±0.0117的输出质量、0.2307±0.0816的重试率以及1.1153±0.0408秒的分配延迟,持续优于其消融配置。DART隔离了持续性和间歇性恶意智能体,实现了99.3%的输出遏制率,并将系统成功率恢复至99.8%。这些结果证明了将声誉、激励、基于DAG的协调以及可验证的区块链赋能治理相结合,以支持自适应和可问责的多智能体协作的潜力。
英文摘要
Large language model (LLM)-based multi-agent systems (MAS) predominantly rely on centralized orchestration and lack formal verification mechanisms for agent reliability, participation, and system-level behavioral alignment. These shortcomings leave open environments severely vulnerable to uncooperative or malicious agents. This work proposes DART, a Directed Acyclic Graph (DAG)-based reputation and incentive regulation framework for trustworthy multi-agent collaboration, combining centralized operational orchestration with blockchain-enabled decentralized governance and accountability. DART unifies DAG workflow orchestration, capability and reputation-aware task allocation, dynamic behavior updates, multi-factor incentives, and smart contract accountability paired with IPFS storage. Under this paradigm, agent selection dynamically balances task alignment, historical reputation, and workload, while post-execution behavioral evidence continuously calibrates agent trust and the probability of future participation. Evaluated across four axes, DART achieves 93.6% Pass@1 on GSM8K and builds a full-stack application in 142 s using two agents, outperforming centralized baselines. Across five independent 150-round longitudinal trials, Full DART achieves a mean task success rate of 93.33 +/- 2.26%, output quality of 0.9357 +/- 0.0117, retry rate of 0.2307 +/- 0.0816, and allocation delay of 1.1153 +/- 0.0408 s, consistently outperforming its ablated configurations DART isolates persistent and intermittent malicious agents, obtaining a 99.3% output containment rate and restoring system success to 99.8%. These results demonstrate the potential of coupling reputation, incentives, DAG-based coordination, and verifiable blockchain-enabled governance to support adaptive and accountable multi-agent collaboration.
发表机构
- Michigan Technological University(密歇根理工大学)
机构由 AI 辅助整理,请以论文原文为准。