发表机构
Aalborg University, Denmark; Bowling Green State University, USA(丹麦奥胡斯大学; 美国布恩威尔州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
探讨大语言模型和人工智能智能体在企业数据集成中面临的挑战,介绍从经典RAG到智能体RAG的演变,研究经济高效集成的优化策略,为构建可靠、可解释和可扩展的数据集成系统指明方向。
AI 中文摘要
大语言模型(LLMs)和人工智能智能体在零样本和少样本设置中的数据集成方面展现出强大潜力。但在企业环境中,由于持续存在的知识差距,它们仍面临显著的准确性和成本挑战。本文设想通过在检索增强生成(RAG)工作流程中运行的基于知识的LLMs和智能体实现可信、可扩展且经济高效的集成。文中追溯了从经典RAG到GraphRAG和KG - RAG(基于知识图谱的RAG)的演变,强调这些范式如何弥合参数知识和上下文知识。在此基础上探索向智能体RAG的转变,其中自主多智能体系统为复杂集成任务自适应地规划、检索、细化和推理。研究了经济高效集成的优化策略,解决大规模企业环境中的计算瓶颈。最后概述了构建可靠、可解释和可扩展的基于知识的集成系统的开放挑战和未来方向。
英文摘要
Large language models (LLMs) and AI agents have demonstrated strong potential for data integration in zero-shot and few-shot settings. However, they continue to face significant accuracy and cost challenges in enterprise environments due to a persistent knowledge gap. This paper envisions trustworthy, scalable, and cost-efficient integration through knowledge-grounded LLMs and agents operating within a retrieval-augmented generation (RAG) workflow. Here, trustworthiness refers to evidence-grounded, verifiable reasoning, where integration decisions are transparently supported by retrieved knowledge, robust against hallucination, and consistent across tasks. We trace the evolution from classic RAG to GraphRAG and KG-RAG (knowledge graph-based RAG), highlighting how these paradigms bridge parametric and contextual knowledge. Building on this trajectory, we explore the shift toward Agentic RAG, where autonomous multi-agent systems adaptively plan, retrieve, refine, and reason for complex integration tasks. We examine optimization strategies for cost-efficient integration, addressing computational bottlenecks in large-scale enterprise settings. Finally, we outline open challenges and future directions toward building reliable, explainable, and scalable knowledge-grounded integration systems.
CommentsTo Appear in the IEEE Data Engineering Bulletin