arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Maat:面向多智能体LLM工作流的独立确定性契约治理

Maat: Independent Deterministic Contract-Based Governance for Multi-Agent LLM Workflows

Uliana Elina

arXiv 2609.34017首次发表:更新:

发表机构

SynWe Group s.r.o.(SynWe集团有限责任公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM多智能体系统中错误传播问题,提出Maat确定性契约治理层,验证交接缺陷,实验显示在多数工作流中提升评分并降低成本,但存在误报,需测试验证器配置。

AI 中文摘要

大语言模型多智能体系统(LLM-MAS)引入了一个特有的可靠性问题:一个智能体产生的错误可能被下游智能体当作上下文接受,并在整个工作流中传播。许多已提出的防护措施依赖于学习型或基于LLM的评判者,而这些评判者的裁决本身是概率性的;我们探究是否可以用一个确定性层来阻止契约可检测的交接缺陷。我们提出了Maat,一个运行时治理层,它根据版本化的工作流契约(即锚点)验证智能体之间的交接,验证或评分路径中不涉及任何语言模型。我们在六个受控领域工作流(6-15个智能体,522次试验)中进行了评估,注入了数据级缺陷,并采用了确定性的七项检查清单。版本1在所有六个工作流中报告了改进(2.9%-26.5%)。发表后审计发现,三个基准评分器将任何提前停止都视为预防了缺陷。在配对试验中,当受治理运行完成或因可归因于已验证缺陷的发现而停止时,五个工作流的评分变化为+7.7%至+29.1%,而在软件开发中保持不变;在可归因的提前停止发生时,模型调用成本下降17%-53%。对所有94次受治理臂停止的人工审查发现35次误报(37%),这些误报由验证器缺陷而非模型行为引起;将这些停止视为失败工作,受治理臂在六个工作流中的四个中得分低于未受治理臂。结果支持对契约可表达缺陷进行确定性交接验证,并表明验证器配置和停止归因本身必须经过测试;这些结果并未确立普遍正确性、幻觉检测或模型无关的有效性。

英文摘要

Large-language-model multi-agent systems (LLM-MAS) introduce a characteristic reliability problem: an error produced by one agent can be accepted as context by downstream agents and propagate across the workflow. Many proposed safeguards rely on learned or LLM-based judges whose verdicts are themselves probabilistic; we ask whether a deterministic layer can instead stop contract-detectable handoff defects. We present Maat, a runtime governance layer that validates agent-to-agent handoffs against a versioned workflow contract, or anchor, with no language model in the validation or scoring path. We evaluate it in six controlled domain workflows (6-15 agents, 522 trials) with injected data-level defects and a deterministic seven-check rubric. Version 1 reported gains in all six workflows (2.9-26.5%). A post-publication audit found that three benchmark scorers credited any early halt as a prevented defect. On paired trials where the governed run completed or halted on a finding attributable to a verified defect, the rubric score changes by +7.7% to +29.1% in five workflows and is flat in software development; model-call cost falls 17-53% where attributable halts occur early. A hand review of all 94 governed-arm halts found 35 false alarms (37%), caused by validator defects rather than model behaviour; counting those halts as failed work, the governed arm scores below the ungoverned arm in four of six workflows. The results support deterministic handoff validation for contract-expressible defects and show that validator configuration and halt attribution must themselves be tested; they do not establish universal correctness, hallucination detection, or model-independent effectiveness.

Comments16 pages, 6 figures, 6 tables. Benchmarks: https://github.com/Lorelys/maat-benchmarks ; CrewAI integration demo: https://github.com/Lorelys/maat-crewai-demo

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑