arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30604cs.CLcs.AI

搜索之后才是难点:在综合、组织与展示知识方面对网络智能体进行基准测试

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge

Alexander Gill, Md Farhan Ishmam, Xuyen Nguyen, Neha Bhat, Parker Henry DeYoung, Fateme Hashemi Chaleshtori, Nathan Stringham, Kenneth Marino, Ana Marasović

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有基准未充分评估智能体助手能力的问题,提出KNOWS基准,通过复杂浏览器任务联合评估检索、综合与展示能力,发现前沿智能体完全成功率不足3%,揭示其端到端助手能力的局限。

中文摘要 AI 辅助

现有的计算机使用智能体基准测试并未充分评估智能体作为助手的能力。一个有用的助手需要在复杂、多步骤的工作流程中检索信息,将其综合成工件(文档、演示文稿、电子表格),并操作程序界面以生成连贯的最终产品。此类工作流程要求推理与综合、复杂任务的分解,以及视觉和空间理解能力。为了研究智能体在这类工作流程上的表现,我们引入了KNOWS,一个基于浏览器的开放式复杂任务基准,该基准联合评估这些能力,每个任务最终都产出一个工件。为了编写任务,我们制定了一套任务设计规范和一个确保任务符合要求的协议。每个任务都配有一个评估器,该程序将确定性检查与LLM判断相结合,以平衡智能体评估中固有的丰富性、可靠性和自动化之间的权衡。我们对前沿的计算机使用智能体和基于浏览器的执行框架进行了评估和分析。它们在部分成功指标上取得了中等分数,但表现最佳的智能体在我们复杂、长时程任务中的完全成功率不足3%。视觉步骤上的失败导致最终工件无法使用,即使智能体完成了超过50%的其他评估步骤。我们的结果揭示了当前智能体作为端到端助手的局限性,并呼吁在工具使用、视觉理解和长时程推理方面取得进展。

英文摘要

Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such workflows demand reasoning and synthesis, decomposition of complex tasks, as well as visual and spatial understanding. To study agents on workflows like these, we introduce KNOWS, a benchmark of open-ended, complex, browser-based tasks that jointly evaluate these capabilities, with each task culminating in a produced artifact. To write tasks, we develop a task design rubric and a protocol for ensuring that tasks meet the requirements. Each task is paired with an evaluator, a program that combines deterministic checks with LLM judgments to balance the richness, reliability, and automation tradeoff inherent to agent evaluation. We evaluate and analyze frontier computer-use agents and browser-based harnesses. They achieve moderate scores on partial-success metrics, but the best performer fully succeeds in fewer than 3% of our complex, long-horizon tasks. Failures on visual steps render the resulting artifacts unusable, even when agents complete more than 50% of other evaluation steps. Our results expose limitations of current agents acting as end-to-end assistants, and call for progress on tool use, visual understanding, and long-horizon reasoning.

发表机构

  • University of Utah(犹他大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑