机场领域对话式AI的RAG方法性能评估
Evaluating the Performance of RAG Methods for Conversational AI in the Airport Domain
- Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
- Royal Schiphol Group(皇家施比尔集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对机场高动态环境下的对话式AI需求,本文对比三种RAG方法的性能,发现Graph RAG准确率最高且幻觉少,SQL RAG幻觉也显著低于传统RAG,二者更适配机场场景。
AI中文摘要:
年旅客量排名前20的机场是高度动态的环境,每日有数千架次航班,且致力于提升自动化程度。为助力这一目标,我们构建了一套对话式AI系统,可让机场工作人员与航班信息系统交互。该系统不仅能解答标准机场问询,还能解析机场术语、行话、缩写以及涉及推理的动态问题。本文构建了三种不同的检索增强生成(Retrieval-Augmented Generation, RAG)方法,包括传统RAG、SQL RAG以及基于知识图谱的RAG(Graph RAG)。实验表明,采用BM25 + GPT-4的传统RAG准确率达84.84%,但偶尔会产生幻觉,这对机场安全存在风险。相比之下,SQL RAG和Graph RAG的准确率分别为80.85%和91.49%,幻觉现象显著减少。此外,Graph RAG在涉及推理的问题上表现尤为突出。基于观察结果,我们推荐在机场环境中优先使用SQL RAG和Graph RAG,因其幻觉更少且具备处理动态问题的能力。
英文摘要:
Airports from the top 20 in terms of annual passengers are highly dynamic environments with thousands of flights daily, and they aim to increase the degree of automation. To contribute to this, we implemented a Conversational AI system that enables staff in an airport to communicate with flight information systems. This system not only answers standard airport queries but also resolves airport terminology, jargon, abbreviations, and dynamic questions involving reasoning. In this paper, we built three different Retrieval-Augmented Generation (RAG) methods, including traditional RAG, SQL RAG, and Knowledge Graph-based RAG (Graph RAG). Experiments showed that traditional RAG achieved 84.84% accuracy using BM25 + GPT-4 but occasionally produced hallucinations, which is risky to airport safety. In contrast, SQL RAG and Graph RAG achieved 80.85% and 91.49% accuracy respectively, with significantly fewer hallucinations. Moreover, Graph RAG was especially effective for questions that involved reasoning. Based on our observations, we thus recommend SQL RAG and Graph RAG are better for airport environments, due to fewer hallucinations and the ability to handle dynamic questions.