arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.11903cs.AI

AEMA:可验证评估框架用于可信和受控的代理LLM系统

AEMA: Verifiable Evaluation Framework for Trustworthy and Controlled Agentic LLM Systems

YenTing Lee, Keerthi Koneru, Zahra Moslemi, Sheethal Kumar, Ramesh Radhakrishnan

首次发表 更新
浏览论文内容

中文总结 AI 辅助

AEMA提出了一种可验证的评估框架,用于评估基于LLM的多代理系统,通过人类监督实现稳定、可追溯的自动化评估。

中文摘要 AI 辅助

评估基于大型语言模型(LLM)的多代理系统仍是一个关键挑战,因为这些系统必须在不断变化的任务中表现出可靠的合作、透明的决策和可验证的性能。现有的评估方法往往局限于单响应评分或狭窄的基准测试,当在多代理规模的企业环境中部署时,缺乏稳定性、可扩展性和自动化。我们提出了AEMA(自适应评估多代理),一个过程感知且可审计的框架,能够在人类监督下计划、执行和汇总跨异构代理工作流的多步骤评估。与单个LLM作为判断者相比,AEMA实现了更高的稳定性、人类对齐性和可追溯的记录,以支持可问责的自动化。我们在使用现实商业场景模拟的企业式代理工作流上的结果表明,AEMA提供了一条透明且可重复的路径,朝着负责任地评估基于LLM的多代理系统迈进。

英文摘要

Evaluating large language model (LLM)-based multi-agent systems remains a critical challenge, as these systems must exhibit reliable coordination, transparent decision-making, and verifiable performance across evolving tasks. Existing evaluation approaches often limit themselves to single-response scoring or narrow benchmarks, which lack stability, extensibility, and automation when deployed in enterprise settings at multi-agent scale. We present AEMA (Adaptive Evaluation Multi-Agent), a process-aware and auditable framework that plans, executes, and aggregates multi-step evaluations across heterogeneous agentic workflows under human oversight. Compared to a single LLM-as-a-Judge, AEMA achieves greater stability, human alignment, and traceable records that support accountable automation. Our results on enterprise-style agent workflows simulated using realistic business scenarios demonstrate that AEMA provides a transparent and reproducible pathway toward responsible evaluation of LLM-based multi-agent systems. Keywords Agentic AI, Multi-Agent Systems, Trustworthy AI, Verifiable Evaluation, Human Oversight

发表机构

  • University of California, San Diego(加州大学圣地亚哥分校)
  • Center for Advanced AI, Accenture(Accenture高级人工智能中心)
  • University of California, Irvine(加州大学伊拉斯姆斯分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑