DI-Bench:系统化生成企业智能体的域内数据智能基准
DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents
AI总结:
针对企业智能体领域基准缺失问题,提出DI-Bench流水线,通过工件关联图生成含731个任务的数据智能基准,评估发现模型在规则修改计算时准确率仅32%。
AI中文摘要:
在特定领域基准上评估企业智能体至关重要,然而公开基准很少评估智能体能否将业务知识与分析计算相结合,且手动构建此类基准成本高昂。我们提出DI-Bench,一个用于生成数据智能(DI)现实基准的流水线,数据智能指从大量企业数据中提取洞察的实践。为模拟需要计算与知识检索相结合的现实DI任务,DI-Bench在数据表、维度、指标和文档之上构建工件关联图,以形成涉及结构化数据及相关知识的问题。真实答案通过查询执行获得,随后进行大语言模型问题生成与验证。将该流水线应用于两个公开数据集,生成了一个包含731个任务的基准,涵盖知识检索、分析计算和基于规则的推理。为展示该基准的区分能力与难度,我们评估了四个模型,揭示了一个重要发现:在检索到的业务规则修改计算的计算任务中,模型准确率仅为32%。
英文摘要:
Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical computation, and constructing such benchmarks manually is costly. We present DI-Bench, a pipeline for generating realistic benchmarks for data intelligence (DI), the practice of extracting insights from large volumes of enterprise data. To emulate realistic DI tasks that require both computation and knowledge retrieval, DI-Bench builds an artifact linkage graph over data tables, dimensions, metrics, and documents to form questions involving structured data and associated knowledge. Ground truth answers are derived via query execution, followed by LLM question generation and validation. Applied to two public datasets, the pipeline produces a 731-task benchmark covering knowledge retrieval, analytical computation, and rule-grounded reasoning. To show the discriminatory capability and difficulty of the benchmark, we evaluate four models, revealing a substantial finding: models achieve only 32% accuracy when doing computational tasks where retrieved business rules modify the computation.