发表机构
King Abdullah University of Science and Technology (KAUST); stc(阿卜杜拉国王科技大学; 沙特电信公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
介绍电信-GAIA双语多模态基准测试,含100个人工验证问答任务,需多跳推理,跨越多种异构源。以沙盒化Docker环境提供,通过规范化精确字符串匹配评分。评估发现其具有挑战性,为企业智能体提供测试平台和构建基准测试的模板。
AI 中文摘要
我们介绍了电信-GAIA,这是一个用于在实际电信运营商数据上评估使用工具的智能体的双语多模态基准测试。它包含100个人工验证的问答任务,有英文和阿拉伯文版本,平均需要4.2跳多跳推理,涉及静态网站快照、合成关系SQL数据库和外部网络存档这三个异构源,涵盖文本、图像和表格模态。基准测试以沙盒化Docker环境提供,通过规范化精确字符串匹配评分,无需大语言模型作为评判即可客观、可重复地评估。评估发现该基准测试具有挑战性,即使最强模型也只能解决71%的任务,在成本预算适中时降至约40%,视觉相关类别最弱,平均后端得分低于30%。电信-GAIA为企业智能体提供了严格、可重复的测试平台和构建封闭领域基准测试的模板。
英文摘要
We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Telco-GAIA comprises 100 human-verified question-answering tasks, in English and Arabic, that each demand multi-hop reasoning (4.2 hops on average) over three heterogeneous sources: a static website snapshot (HTML, images, and linked PDFs), a synthetic relational SQL database, and external web archives, spanning text, image, and tabular modalities. The benchmark is delivered as a sandboxed Docker environment and scored by normalized exact string matching, making evaluation objective, deterministic, and reproducible over time without any LLM-as-a-Judge. Evaluating a purpose-built reference agent across twelve commercial and open LLMs, we find Telco-GAIA challenging: even the strongest model solves only 71% of tasks; under a moderate cost budget, this falls to about 40%, and the visually grounded categories remain the weakest, where the average backend scores below 30%, leaving substantial headroom in document and image understanding. Telco-GAIA offers a rigorous, reproducible testbed for enterprise agents and a template for constructing closed-domain benchmarks.