arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HyperBrowseComp:面向网页浏览代理的多语言与多模态压力测试

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

Alham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri, Seung Hun Eddie Han, Ryandito Diandaru, Qinrong Cui, Jan Christian Blaise Cruz, Badrinath Chandana, Peerawat Chomphooyod, Ahmed Attia, Jonibek Mansurov, Emilio Villa-Cueva, Canh Duong Nguyen, Imran Turganov, Minghao Wu, Peerat Limkonchotiwat, Irina Nikishina

arXiv 2610.03574首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence; Mila – Quebec Artificial Intelligence Institute; Inception AI; Alibaba Group; AI Singapore(穆罕默德·本·扎耶德人工智能大学; Mila – 魁北克人工智能研究所; Inception AI; 阿里巴巴集团; AI新加坡)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HyperBrowseComp是一个包含423个多语言多模态难题的网页浏览代理基准,通过过滤简单问题并评估多个模型,为跨语言证据搜索提供挑战性测试。

AI 中文摘要

我们推出了HyperBrowseComp,这是一个多语言和多模态的浏览基准测试,包含423个由母语者或高水平使用者编写并经过人工验证的问题,覆盖13种语言。这些问题设计得极具挑战性。每个问题都针对一个简洁、可公开验证的答案,而找到该答案需要定位晦涩的证据、遵循多步线索链,或检查异构来源,如视频、扫描文档、图像或地图。较简单的问题通过使用无互联网访问的模型进行评估而被过滤掉,以降低仅凭参数化知识即可回答的可能性。我们使用提供商原生的搜索和共享的外部检索工具,在统一的代理协议下评估了多个模型。为了将模型性能与努力程度置于背景中,我们还对部分问题进行了人工评估。HyperBrowseComp为跨语言和证据模态的持久信息搜索提供了一个具有挑战性的测试平台,其难度源于在开放网络上发现并连接证据。

英文摘要

We introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers. Questions are designed to be extremely challenging. Each question targets a concise, publicly verifiable answer whose discovery requires locating obscure evidence, following multi-step clue chains, or inspecting heterogeneous sources such as videos, scanned documents, images, or maps. Easier questions are filtered out by evaluating them with models without internet access to reduce the likelihood that they can be answered with parametric knowledge alone. We evaluate several models using provider-native search and a shared external retrieval harness under a common agent protocol. To contextualize model performance and effort, we also conduct a human evaluation on a sample of the questions. HyperBrowseComp provides a challenging testbed for persistent information seeking across languages and evidence modalities, with difficulty arising from discovering and connecting evidence on the open web.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑