arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11584cs.AI

EnterpriseRAG:非理想企业检索场景下LLM指令遵循与鲁棒性的基准测试

EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval

  • Jiutian Research, China Mobile(中国移动九天研究院)

机构由 AI 辅助整理,请以论文原文为准。

Huiqi Miao, Xinbao Sun, Bo Wang, Fanyu Meng, Lijun Mei, Na Wu, Di Jin, Chao Deng, Junlan Feng

中文总结 AI 辅助

研究针对企业RAG部署的可靠性问题,推出含983个样本的EnterpriseRAG基准,评估13个LLM发现其指令遵循崩溃,为生产级RAG提供可复现的评估基础

中文摘要 AI 辅助

企业RAG部署面临关键的可靠性差距:虽然大型语言模型(LLM)满足80%的单个约束条件,但仅有26.8%的响应同时满足所有要求,显示出57分的编排差距。现有基准假设检索干净且查询简单,无法捕捉存在噪声文档和多维约束共存的生产环境。我们推出EnterpriseRAG,这是一个涵盖6个领域、经983个专家验证样本的基准,系统模拟了现有工作中缺失的三种失败模式:检索噪声、知识缺口和事实冲突,同时结合复杂指令。对13个最先进的LLM的评估显示存在严重的指令遵循崩溃,即高单约束满意度掩盖了低整体合规性。关键发现揭示了在知识缺口和事实冲突下的深层障碍,即便采用推理增强推理也是如此,这表明生产级RAG需要明确的上下文感知协议和校准判断。EnterpriseRAG为衡量和缩小这些差距提供了可复现的基础,直接为企业级RAG系统的部署决策提供信息,我们将在发表后发布该基准和评估框架。

英文摘要

Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retrieval with simple queries, failing to capture production conditions where noisy documents and multi-dimensional constraints coexist. We introduce EnterpriseRAG, a benchmark of 983 expert-validated samples across six domains that systematically simulates three failure modes absent from prior work: retrieval noise, knowledge gaps, and factual conflicts, coupled with complex instructions. Evaluation of 13 state-of-the-art LLMs reveals a severe instruction adherence collapse, where high per-constraint satisfaction masks low holistic compliance. Critical findings expose deep barriers under knowledge gaps and factual conflicts, even with reasoning-enhanced inference, indicating production RAG requires explicit context-aware protocols and calibrated judgment. EnterpriseRAG provides a reproducible foundation for measuring and closing these gaps, directly informing deployment decisions for enterprise-scale RAG systems. We will release the benchmark and evaluation framework upon publication.

↑