BI-Agent 与 BI-Bench:迈向端到端商业智能的自动化
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
- UIUC(伊利诺伊大学厄巴纳-香槟分校)
- Microsoft Research(微软研究院)
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出 BI-Agent 与 BI-Bench,首个系统评估 LLM 端到端 BI 能力的基准,通过工具增强与后训练显著提升准确率。
AI中文摘要:
商业智能(BI)是企业决策的基石,被 Power BI 和 Tableau 等软件中的企业用户广泛使用。在传统的 BI 工作流程中,用户需要先准备数据,包括(1)识别相关表格,(2)执行数据转换,(3)构建连接关系,然后才能(4)回答其业务问题。这些步骤可能复杂且耗时,使得 BI 具有挑战性。鉴于大型语言模型(LLM)在处理数据方面的强大能力,我们研究了它们端到端回答 BI 问题的能力,而无需用户手动执行繁琐的准备步骤。为此,我们从公共来源收集了大量真实的 BI 项目,并手动从真实用户仪表板中提取(问题,真实答案)对。由此产生的基准 BI-Bench 是第一个系统研究 LLM 在端到端 BI 上能力的基准。我们发现,即使是前沿 LLM 在 BI-Bench 上表现也很差,准确率低于 50%。为了解决这些局限性,我们设计了一个工具增强的 BI-Agent,它将 BI 工作流分解为结构化数据上的子任务,例如搜索、连接和转换,并在 BI 各阶段协调专门的数据管理方法。此外,我们开发了一个后训练框架,从真实 BI 项目中合成训练轨迹,使 BI-Agent 能够使用监督微调(SFT)和强化学习(RL)进一步后训练。BI-Agent 在使用普通 LLM 时实现了高达 40 个百分点的显著准确率提升,而后训练的 BI-Agent 则带来了高达 30 个百分点的提升。我们的结果强调了在复杂 BI 工作流中将工具增强推理与领域特定后训练相结合的重要性,并为未来的研究指出了有前景的方向。
英文摘要:
Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau. In traditional BI workflows, users need to prepare data by (1) identifying relevant tables, (2) performing data transformations, and (3) building join relationships, before they can (4) answer their business questions. These steps can be complex and time-consuming, making BI challenging. Given the strong capabilities of large language models (LLMs) in working with data, we study their ability to answer BI questions end-to-end, without requiring users to manually perform the tedious preparation steps. To do this, we harvest a large collection of real-world BI projects from public sources, and manually extract pairs of (questions, ground-truth answers) from real user dashboards. The resulting benchmark, BI-Bench, is the first benchmark to systematically study LLMs' ability on end-to-end BI. We find that even frontier LLMs perform poorly on BI-Bench, with less than 50% accuracy. To address their limitations, we design a tool-augmented BI-Agent that decomposes BI workflows into subtasks on structured data, such as search, join, and transform, and orchestrates specialized data management methods across BI stages. Furthermore, we develop a post-training framework that synthesizes training trajectories from real BI projects, enabling BI-Agent to be further post-trained using both supervised fine-tuning (SFT) and reinforcement learning (RL). BI-Agent achieves substantial accuracy gains of up to 40 percentage points with vanilla LLMs, and post-trained BI-Agent yields gains of up to 30 points. Our results highlight the importance of combining tool-augmented reasoning with domain-specific post-training in complex BI workflows, and point to promising directions for future research.