发表机构
Epiq AI Labs; Cornell University(Epiq AI实验室; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对企业级LLM评估的数据集难题,构建了含23万+文档的CorporateBench基准,从信息提取与知识库查询维度评估5种LLM,发现输入规模接近实际时性能下降,填补了相关评估指标空白。
AI 中文摘要
大型语言模型(LLMs)越来越能够回答关于企业级文档集合的复杂问题,但评估工作面临困难:企业不愿共享内部通信内容,而合成数据集又过于简单。我们提出CorporateBench(CB),这是一个经人工验证的多任务问答基准,其规模接近LLMs在企业通信网络中遇到的实际条件,评估语料库超过23万份文档。CB通过4家员工规模从12到10000人的合成生成企业,从两个维度(信息提取和知识库查询)对LLMs进行评估。每个语料库均来自描述一致世界的时序演化知识库,确保即使在数十万份文档间也具备跨文档逻辑一致性。我们在CB上评估了5种LLMs,发现随着输入规模接近实际尺度,性能愈发不佳。CB为LLMs开发者提供了企业通信推理的评估指标,填补了基准测试生态系统中的关键空白。
英文摘要
LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, with evaluation corpora surpassing 230,000 documents. CB evaluates LLMs across two dimensions (information extraction and knowledge base querying) through four synthetically generated firms ranging from 12 to 10,000 employees. Each corpus is sampled from a temporally evolving knowledge base describing a consistent world, guaranteeing cross-document logical consistency even across hundreds of thousands of documents. We evaluate five LLMs on CB, revealing increasingly poor performance as input size approaches realistic scales. CB provides LLM developers a metric for corporate communication reasoning, filling a crucial gap in the benchmarking ecosystem.
CommentsAccepted to EMNLP Findings