从朴素检索增强生成到深度智能检索:用于合规监管的不断演进的上下文工程管道
From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance
浏览论文内容
中文总结 AI 辅助
研究针对RAG朴素实现的局限,追溯安大略发电公司检索管道演变,历经多阶段形成PEA-CAE架构,表明上下文工程对监管语料库更具优势,深度智能检索进展反映经典思想,迭代证据获取等渐成企业问答基础。
中文摘要 AI 辅助
检索增强生成(RAG)是将大语言模型应用于企业文档语料库的主导范式,但朴素实现随着语料库规模和查询复杂度增加面临硬限制。本文追溯了安大略发电公司(OPG)用于合规监管和费率案例分析的生产检索管道的演变。研究了朴素RAG、带重新排序的混合检索、智能函数调用检索以及具有基于代码的工具合成和显式规划的深度多智能体架构等阶段,识别促使每次转变的失败模式和权衡。将成熟架构形式化为具有成本感知升级的渐进证据获取(PEA-CAE)。研究结果表明,对于大型、不断演变的监管语料库,上下文工程比特定领域微调更易处理且经济可行。更广泛地说,向深度智能检索的进展反映了经典信息检索思想,并引入了自适应查询重新制定、渐进文档发现和分层子智能体总结等实用系统原语。操作跟踪进一步支持现代检索系统基于搜索的性质,其中迭代证据获取和自适应规划越来越多地取代单遍检索成为企业级问答的基础。
英文摘要
Retrieval-augmented generation (RAG) is the dominant paradigm for applying large language models (LLMs) to enterprise document corpora, yet naive implementations encounter hard limits as corpus scale and query complexity grow. This paper traces the evolution of a production retrieval pipeline at Ontario Power Generation (OPG) for regulatory compliance and rate case analysis under Ontario Energy Board (OEB) reporting requirements. We examine successive stages: naive RAG, hybrid retrieval with re-ranking, agentic function-calling retrieval, and a deep multi-agent architecture with code-based tool synthesis and explicit planning, and identify the failure modes and tradeoffs that motivated each transition. We formalize the mature architecture as Progressive Evidence Acquisition with Cost-Aware Escalation (PEA-CAE): begin with low-cost, high-precision retrieval and escalate to full-document reads only when the expected evidence gain justifies latency and cost. Our findings show that context engineering is a more tractable and economically viable path than domain-specific fine-tuning for large, evolving regulatory corpora. More broadly, the progression toward deep agentic retrieval mirrors classical information retrieval ideas, introducing adaptive query reformulation, progressive document discovery, and hierarchical subagent summarization as practical system primitives. Operational traces further support the search-based nature of modern retrieval systems, where iterative evidence acquisition and adaptive planning increasingly replace single-pass retrieval as the foundation for enterprise-scale question answering.
发表机构
- Ontario Power Generation(安大略发电公司)
- Dalhousie University(达尔豪斯大学)
机构由 AI 辅助整理,请以论文原文为准。