arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25765cs.CLcs.DB

WorkSurface-Bench:在多表面知识路由上对企业代理进行基准测试

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing

Hao Liang, Meiyi Qiang, Sizhe Qiu, Linzhuang Sun, Wentao Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对企业代理整合异构知识源时表面路由能力评估不足的问题,提出WorkSurface-Bench基准测试,含多种任务且答案可审计,评估多个模型主干,揭示表面选择与任务完成关系及干预效果,并发布相关资源。

中文摘要 AI 辅助

企业代理常常需要整合异构知识源,如文档、表格和依赖图。现有基准测试通常不区分代理是否先选择合适的知识源来评估检索或工具使用。我们引入WorkSurface-Bench作为评估此能力的基准测试。它包含1151个原子任务,答案可审计。我们在六种控制代理设置下评估了四个模型主干,结果表明正确的表面选择对任务完成必要但不充分。匹配干预显示表面提示和去除无关工具的效果。最后数据集等已发布。

英文摘要

Enterprise agents often need to integrate heterogeneous knowledge sources: documents for narrative facts, tables for computation, and dependency graphs for file relationships. Existing benchmarks typically evaluate retrieval or tool use without distinguishing whether an agent first selects the appropriate knowledge sources. We introduce WorkSurface-Bench, a benchmark for evaluating this capability as surface routing. It contains 1,151 atomic tasks derived from persona-scoped Workspace-Bench-Lite workspaces, spanning document, table, graph, and cross-surface questions. Its reference answers are auditable: table answers are reproduced through executed DuckDB queries, document answers are grounded in verified text spans, and graph answers are traced to source dependency annotations. We evaluate four model backbones across six controlled agent settings, yielding 27,624 protocol-error-free trajectories. Under gold-constrained tool access, agents achieve 98.7-99.8 Route F1, while Answer remains only 56.1-75.3 percent, showing that correct surface selection is necessary but insufficient for task completion. Matched interventions further show that surface hints improve Answer for three of four models, whereas removing irrelevant tools primarily improves routing and efficiency. In an independent three-annotator audit, all 200 sampled tasks pass all six quality criteria by majority vote, with 192 receiving unanimous judgments on every criterion. We release the dataset, construction pipeline, scoring code, and agent harness at https://github.com/haolpku/WorkSurface-Bench.

发表机构

  • Peking University(北京大学)
  • Zhongguancun Academy(中关村科学城)
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

↑