arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13856cs.CL

ShopEase:基于生成式AI的多智能体框架,利用混合检索增强生成实现智能企业客户支持

ShopEase: A Generative AI-Based Multi-Agent Framework for Intelligent Enterprise Customer Support Using Hybrid Retrieval-Augmented Generation

发表机构印度理工学院卡哈拉格普尔分校
查看机构详情
  • Indian Institute of Technology Kharagpur(印度理工学院卡哈拉格普尔分校)

机构由 AI 辅助整理,请以论文原文为准。

Aakash Kumar Tiwari, Somesh Kumar

首次发表
浏览论文内容

中文总结 AI 辅助

ShopEase提出基于生成式AI的多智能体框架,结合FAISS与BM25混合检索及LLaMA 3.2,在2632个查询上仅FAISS达85.37%准确率,优于重排序配置,实现高效企业客户支持。

中文摘要 AI 辅助

企业客户支持系统必须正确回答客户问题、检索正确的政策信息、利用客户上下文,并在需要时将困难案例转交给人工客服。本文介绍了ShopEase,一个基于生成式AI的多智能体框架,用于企业客户支持。该系统结合了六个组件:意图识别、客户关系管理、记忆、混合RAG、升级处理和监督,并使用通过Ollama在本地运行的LLaMA 3.2进行响应生成。检索模块结合了FAISS(稠密检索)和BM25(稀疏检索),并评估了六种配置:仅BM25、仅FAISS、公平RRF、加权RRF、带交叉编码器的RRF以及带交叉编码器的Top-10混合。策略类别不是通过意图与政策之间的固定映射来确定,而是直接从检索到的文档中决定。该系统在2632个保留的客户查询上进行了评估,涵盖六个类别:退款、退货、运输、取消、损坏产品和未知。仅FAISS实现了最高准确率85.37%(2247个正确预测),紧随其后的是加权RRF,准确率为85.07%。仅BM25的准确率仅为55.74%。添加交叉编码器重排序并未改善结果:带交叉编码器的RRF达到83.24%,带交叉编码器的Top-10混合达到81.88%,同时还增加了响应延迟。类别层面的分析显示,在运输、取消和退货方面表现强劲,而未知查询仍然是错误的主要来源。使用McNemar检验进行的统计测试显示,仅FAISS和加权RRF之间没有显著差异,尽管两者都显著优于公平RRF和交叉编码器配置。总体而言,稠密检索在该数据集上提供了最佳准确率,而额外的重排序增加了处理时间,却没有改善分类性能。

英文摘要

Enterprise customer support systems must answer customer questions correctly, retrieve the right policy information, use customer context, and pass difficult cases to human agents when needed. This paper presents ShopEase, a Generative AI-based multi-agent framework for enterprise customer support. The system combines six components: Intent, CRM, Memory, Hybrid RAG, Escalation, and Supervisor, and uses LLaMA 3.2 running locally through Ollama for response generation. The retrieval module combines FAISS (dense retrieval) and BM25 (sparse retrieval), and six configurations are evaluated: BM25-only, FAISS-only, Fair RRF, Weighted RRF, RRF with Cross-Encoder, and Top-10 Hybrid with Cross-Encoder. Instead of using a fixed mapping between intent and policy, the policy category is decided directly from the retrieved documents. The system was evaluated on 2632 held-out customer queries across six categories: Refund, Return, Shipping, Cancellation, Damaged Product, and Unknown. FAISS-only achieved the highest accuracy of 85.37\% (2247 correct predictions), closely followed by Weighted RRF at 85.07\%. BM25-only achieved only 55.74\% accuracy. Adding cross-encoder reranking did not improve results: RRF with Cross-Encoder reached 83.24\%, and Top-10 Hybrid with Cross-Encoder reached 81.88\%, while also increasing response latency. Category-level analysis shows strong performance on Shipping, Cancellation, and Return, while Unknown queries remain the main source of errors. Statistical testing using McNemar's test shows no significant difference between FAISS-only and Weighted RRF, though both perform significantly better than Fair RRF and the cross-encoder configurations. Overall, dense retrieval gives the best accuracy on this dataset, and additional reranking adds processing time without improving classification performance.

↑