AI 中文总结
该研究针对企业智能体查询结构化数据的架构挑战,划分七个维度提出设计框架,经实验验证其可消除授权违规,还给出评估协议与开放问题。
AI 中文摘要
检索增强生成(RAG)已成为连接大型语言模型与企业知识的常用架构。大多数RAG系统会检索非结构化文档(PDF、维基页面、支持工单),并将其输入LLM以进行摘要生成或问答。然而,越来越多的企业智能体必须查询结构化数据:关系型数据库、数据仓库和分析API,其中答案是计算结果而非检索到的段落。结构化数据查询会带来文档RAG管道无需做出的决策,我们将这些决策分为七个维度:检索语义、授权、意图识别、实体解析、评估、故障模式和延迟。针对每个维度,我们描述了基线假设,解释其在结构化数据场景中的局限性,并阐述通用架构模式。作为支撑证据,一项受控合成研究表明,基于该框架构建的分阶段智能体,在受控条件下消除了直接翻译执行基线的授权违规问题。主要成果是一个面向设计的框架、一套评估协议,以及受管结构化数据智能体的一系列开放问题。
英文摘要
Retrieval-augmented generation (RAG) has become a common architecture for connecting large language models to enterprise knowledge. Most RAG systems retrieve unstructured documents (PDFs, wiki pages, support tickets) and feed them to an LLM for summarization or question answering. A growing class of enterprise agents, however, must query structured data: relational databases, data warehouses, and analytics APIs where the answer is a computed result, not a retrieved passage. Structured-data querying forces decisions that a document-RAG pipeline never has to make. We group them into seven dimensions: retrieval semantics, authorization, intent recognition, entity resolution, evaluation, failure modes, and latency. For each dimension, we characterize the baseline assumption, explain its limitation for structured data, and describe a generic architectural pattern. As supporting evidence, a controlled synthetic study shows that a staged agent built on this framework eliminates the authorization violations of a direct translate-and-execute baseline under controlled conditions. The primary result is a design-oriented framework, an evaluation protocol, and a set of open problems for governed structured-data agents.
Comments9 pages, 4 tables, 2 figures