arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23806cs.AI

WorkWorlds:用于评估智能体在工作场所任务中表现的基础设施

WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks

Yining Hua, Levi Lian

首次发表
浏览论文内容

中文总结 AI 辅助

WorkWorlds通过分离组织状态与任务规范,在合成制药公司中评估智能体,发现从任务策划上下文到完整工作场所,证据访问和标准通过率显著下降,差异主要发生在证据获取前。

中文摘要 AI 辅助

许多知识工作基准测试是围绕单个任务构建的,每个任务所需的上下文是在任务被指定时或之后与任务一起选择的。这种设计在为一个任务组装的环境中衡量类似工作场所任务的表现。当任务规范指导选择哪些上下文时,评估会将任务信息编码到环境中,并预先完成部分通常需要工作场所表现所需的信息定位工作。我们引入了WorkWorlds,一种将组织状态与任务规范分离的评估基础设施。一个世界首先固定一个修订版本、日期和员工座位,并物化该员工可以访问的组织状态;任务仅在之后引入。我们在一个主要的合成制药公司中实现了WorkWorlds,包含6个员工座位上的8个测量任务,并构建了额外的组织世界。在192个匹配的评估中,从任务策划的上下文转移到完整的角色可见工作场所,证据访问率从90.4%降至74.5%,标准通过率从79.4%降至68.2%,而证据访问条件下的通过率几乎保持不变;测量的差异大部分发生在智能体达到足够证据之前。

英文摘要

Many knowledge-work benchmarks are constructed around individual tasks, with the context needed for each task selected together with or after the task has been specified. This design measures performance on workplace-like tasks in an environment assembled for the task. When task specification guides which context is selected, the evaluation can encode task information into the environment and pre-complete part of the information-localization work that workplace performance normally requires. We introduce WorkWorlds, an evaluation infrastructure that separates organizational state from task specification. A world first fixes a revision, date, and employee seat and materializes the organizational state that employee can access; tasks are introduced only afterward. We implement WorkWorlds in a primary synthetic pharmaceutical company with 8 measured tasks across 6 employee seats, and construct additional organizational worlds. Across 192 matched evaluations, task-level curation increased evidence access by 17.6 percentage points, from 72.8% to 90.4%, and criterion pass by 8.7 points, from 68.0% to 76.7%, while pass conditional on evidence access remained nearly unchanged; most of the measured difference occurred before the agent reached sufficient evidence.

发表机构

  • Harvard University(哈佛大学)
  • Agent Evaluation Science Inc.(Agent Evaluation Science 公司)
  • Raycaster
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑