用于自动化企业分析与洞察生成的多智能体平台
A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation
浏览论文内容
中文总结 AI 辅助
本文提出基于CrewAI的多智能体框架,含五个专业智能体,经300个测试用例验证,功能准确率95.3%,较单智能体基线提升显著,可用于自动化企业分析与洞察生成。
中文摘要 AI 辅助
本文提出了一个基于CrewAI的多智能体框架,用于对话式商业智能。五个专业AI智能体按顺序流水线运作,处理自然语言查询、检索并分析数据、通过模型上下文协议(MCP)生成可视化内容,并交付可操作的洞察。该平台具备纵深防御安全架构,支持多租户数据隔离,还拥有查询参数化机制,可将对话式洞察转换为可复用的仪表板组件。对涵盖合成及生产企业数据集的300个端到端测试用例进行评估,结果显示其功能准确率达95.3%,平均响应延迟为24秒,经LLM-as-a-Judge框架评估的响应质量评分为4.52/5.0,无幻觉率为93.0%;相较于单智能体基线,准确率提升了22.6个百分点,质量提升了20.2%。跨四个大语言模型后端的评估及人类专家验证,确认了该架构的可泛化性与评估者可靠性。消融研究证实,数据分析与报告聚合智能体是输出质量的主要驱动因素。
英文摘要
This paper proposes a multi-agent framework built on CrewAI [1] for conversational business intelligence. Five specialized AI agents operate in a sequential pipeline to process natural language queries, retrieve and analyze data, generate visualizations via the Model Context Protocol (MCP) [2], and deliver actionable insights. The platform features a defense-in-depth security architecture for multi-tenant data isolation and a query parameterization mechanism for transforming conversational insights into reusable dashboard components. Evaluation across 300 end-to-end test cases spanning synthetic and production enterprise datasets demonstrates 95.3% functional accuracy, a mean response latency of 24 seconds, and a response quality score of 4.52/5.0 as assessed by an LLM-as-a-Judge framework, with a 93.0% hallucination-free rate, representing a 22.6 percentage point accuracy improvement and 20.2% quality gain over a single-agent baseline. Cross-model evaluation across four LLM backends and human expert validation confirm architectural generalizability and evaluator reliability. An ablation study confirms that the Data Analysis and Report Aggregation agents are the primary drivers of output quality.
发表机构
- Rakuten India Enterprise Private Limited(乐天印度企业私人有限公司)
机构由 AI 辅助整理,请以论文原文为准。