arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BatchDAG:用于企业数据上可扩展即席分析的大语言模型规划执行图

BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data

Anupreet Walia

arXiv 2607.18241首次发表:更新:

AI 中文总结

研究针对企业数据跨实体分析问题,提出BatchDAG系统,大语言模型生成操作DAG,经确定性引擎评估。通过实体感知批处理优化,减少大语言模型调用。实验显示其质量高、溯源性好,能高效处理查询,是通用编排层替代手工工作流程。

AI 中文摘要

大语言模型(LLMs)在分析单个文档方面表现出色,但在处理企业规模数据集上的详尽、跨实体分析问题时会因上下文溢出、实体属性丢失以及顺序工具调用的线性延迟而失效。我们提出了BatchDAG系统,其中大语言模型生成操作的类型化有向无环图(DAG)——SQL查询、语义搜索、内存转换、并行扇出和单次分析——由确定性引擎通过拓扑波并行和结构化JSON数据流进行评估。关键优化是实体感知批处理,在扇出前按逻辑实体对行进行分组,将大语言模型调用减少多达47倍。BatchDAG主要不是在准确性上优于手工优化的管道;相反,它是一个通用编排层,用一个从自然语言生成适当执行策略的单一系统取代多个手工设计的工作流程。在对12个转录密集型查询的控制实验中,BatchDAG(3.74/5)实现了与专家设计的管道(3.25/5)相当的质量,并且显著优于ReAct代理(3.09/5,p<0.01),具有更高的溯源性(77%的转录证据率,而基线为46 - 60%)。控制消融表明,与散文摘要相比,结构化JSON中间件将幻觉减少了27%(配对t检验,p = 0.107,n = 12)。规划器在300次规划调用中实现了98.8%的有效DAG率。在该网址的生产环境中,BatchDAG在不到60秒的时间内处理超过50,000次会议的查询,按照已公布的GPT - 5.1定价计算,每次查询的成本为0.02 - 0.24美元。

英文摘要

Large language models (LLMs) excel at analyzing individual documents but break down on exhaustive, cross-entity analytical questions over enterprise-scale datasets due to context overflow, loss of per-entity attribution, and linear latency from sequential tool calls. We present BatchDAG, a system in which an LLM generates a typed directed acyclic graph (DAG) of operations -- SQL queries, semantic searches, in-memory transforms, parallel fan-outs, and single-shot analyses -- which a deterministic engine evaluates with topological-wave parallelism and structured JSON data flow. A key optimization, entity-aware batching, groups rows by logical entity before fan-out, reducing LLM calls by up to 47x. BatchDAG is not primarily an accuracy improvement over hand-optimized pipelines; rather, it is a general-purpose orchestration layer that replaces multiple hand-engineered workflows with a single system that generates the appropriate execution strategy from natural language. In controlled experiments on 12 transcript-heavy queries, BatchDAG (3.74/5) achieves quality comparable to an expert-designed pipeline (3.25/5) and significantly outperforms a ReAct agent (3.09/5, p<0.01), with superior provenance (77% transcript evidence rate vs. 46-60% for baselines). A controlled ablation shows structured JSON intermediates reduce hallucinations by 27% versus prose summaries (paired t-test, p=0.107, n=12). The planner achieves 98.8% valid-DAG rate across 300 planning calls. In production at Brevian.ai, BatchDAG processes queries over 50,000+ meetings in under 60 seconds, with measured per-query costs of $0.02-$0.24 at published GPT-5.1 pricing.

Comments9 pages, 2 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑