WinSyn:用于真实企业问答评估的自动化流水线
WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation
- Microsoft Research(微软研究院)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对现有基准缺乏企业复杂性的问题,提出WinSyn自动化流水线,生成模拟多角色、多月的合成邮件及问答数据,评估显示前沿模型得分低于80%,凸显真实高复杂度评估数据的重要性。
中文摘要 AI 辅助
企业环境为问答智能体提供了一个具有挑战性的环境,这些智能体通常依赖检索增强生成、深度研究(DR)及相关技术。这一挑战很大程度上源于企业数据的复杂性:信息通常分散在不断演变且可能相互矛盾的电子邮件、聊天消息、文档和其他工件中。现有基准通常具有有限的实际复杂性、短形式的回答和不自然的查询,因此往往无法捕捉企业环境中的挑战。在这项工作中,我们引入了一个自动化流水线,用于生成反映现实工作场景的合成电子邮件数据集,以及基于数据的长短形式问题和黄金答案。我们的方法模拟了持续数月、涉及多达25名跨多个角色互动的员工的长运行企业项目。数据强调模糊性、信息分布性和自然发生的查询。为验证该流水线,我们使用最新的前沿模型在数据集上评估了几个标准的智能体基线。我们发现,每个数据集上所有查询的平均聚合分数仍低于80%,表明还有显著的改进空间。这些发现表明,企业部署仍需更多工作,并强调了现实、高复杂性评估数据对于开发更强大的现实世界企业DR系统的重要性。
英文摘要
Enterprise settings provide a challenging environment for question-answering agents, which often rely on Retrieval-Augmented Generation, Deep Research (DR), and related techniques. Much of this challenge comes from the complexity of enterprise data: information is often spread across evolving and potentially conflict- ing emails, chat messages, documents, and other artifacts. Existing benchmarks typically have limited real-world complexity, short-form responses, and unnatural queries, so they often fail to capture the challenges of enterprise settings. In this work, we introduce an automated pipeline for generating synthetic datasets of emails reflecting realistic workplace scenarios, along with long- and short-form questions and gold answers grounded in the data. Our method simulates long-running enterprise projects spanning several months and involving up to 25 interacting employees across multiple roles. The data emphasizes ambiguity, distributed information, and naturally occurring queries. To validate the pipeline, we evaluate few standard agentic baselines on our datasets using the latest frontier models. We find that aggregate scores averaged over all queries remain below 80% for each dataset, indicating significant room for improvement. These findings suggest that more work remains to be done for enterprise deployment and underscore the importance of realistic, high-complexity evaluation data for developing stronger real-world enterprise DR systems.