arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmegaUse-OfficeVal:基于经济基准的长周期办公套件任务LLM智能体评测基准

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Jingbo Zhou, Yusai Zhao, Qi Bao, Jingjia Cao, Zhenghai Chen, Chang Gao, Kaiqi Guo, Muxin Guo, Mingxuan Li, Xinjiang Lu, Yanru Ma, Yixiong Xiao, Zenghui Zhang, Le Zhang, Hua Wu

arXiv 2607.27155首次发表:更新:

发表机构

Baidu Inc.(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OmegaUse-OfficeVal是带经济基准的长周期办公套件任务LLM智能体评测基准,含100项任务,评测发现前沿LLM成本低速度快但质量未达人类水平,代码与数据集已开源。

AI 中文摘要

大型语言模型(LLM)智能体被期望越来越多地协助用户完成任务,但现有基准在评估智能体能否以合理成本执行办公套件工作流程方面支持有限。我们引入OmegaUse-OfficeVal,这是一个用于评测长周期办公套件任务LLM智能体的基准,具备任务级经济基准。该基准包含100项任务,这些任务源自从业者提出的办公套件需求,并通过隐私保护流程调整而来。平均而言,这些任务需要2.32小时的人工劳动完成。该基准的一个重要特征是,每项任务都配有两个经济信号:人工劳动时间和任务价格代理,这些信号可实现人工成本与LLM推理成本的直接比较,以及价值加权评估。为支持稳定评测,我们基于细粒度规则开发了基于代码的验证器。我们对多个前沿LLM以及人类基线进行了评测,尽管所有被评测的LLM都比人类工作者便宜得多且速度快得多,但它们尚未达到人类水平的交付质量。代码和数据集已完全开源,更多信息可在我们的项目网站获取:this https URL。

英文摘要

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality. The code and dataset are fully open-sourced, and more information is available on our project website: https://omegause-officeval.github.io.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑