发表机构
University of Houston; Alexandria University; Bogazici University; Ankara Yıldırım Beyazıt University; ESI, Algiers; University of Algiers 1; York St John University London; Air University, Islamabad; Mohammed Premier University; Universidad Católica San Pablo; Busitema University; University of Laghouat; University of Tunis El Manar; Tribhuvan University; IIT Roorkee; Shobhit Institute of Engineering and Technology; MIT World Peace University; University of Kinshasa; University of Bertoua; University of the Witwatersrand(休斯顿大学; 亚历山大大学; 博阿齐奇大学; 安卡拉耶尔德勒姆贝亚泽特大学; 阿尔及尔ESI; 阿尔及尔第一大学; 约克圣约翰大学伦敦校区; 伊斯兰堡空军大学; 穆罕默德第一大学; 圣巴勃罗天主教大学; 布西特马大学; 拉格瓦特大学; 突尼斯埃尔马纳尔大学; 特里布万大学; 印度理工学院鲁尔基分校; 肖比特工程技术学院; MIT世界和平大学; 金沙萨大学; 贝尔图阿大学; 威特沃特斯兰德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLMs在端到端GNN编码任务上能力评估缺失的问题,提出首个竞赛基准GNN-CB,含18个竞赛,统一自动化评估,发现LLMs难及人类顶级表现。
AI 中文摘要
大型语言模型(LLMs)在编码和推理基准测试中展现出强大性能;然而,它们解决图结构机器学习问题的能力在很大程度上仍未得到探索。特别是,目前尚无基准测试评估LLMs能否在现实竞赛设置下自主解决端到端的图神经网络(GNN)编码任务。为弥补这一空白,本文介绍了GNN-CB,这是首个基于竞赛的基准,用于评估人类和LLMs在GNN编码任务上的表现。GNN-CB包含18个精选竞赛,涵盖节点级、边级和图级预测,涉及多种图类别、领域和难度层级。所有提交均通过统一的自动化流程进行评估,该流程包含隐藏测试集和标准化评分。人类参与者在受控竞赛约束下解决问题,而LLMs则通过基于计划-然后-编码范式的冻结零样本提示协议进行评估,并带有有界的执行-修复循环。该基准还支持在同一协议内进行非智能体和基于自主智能体的评估。在我们评估的协议下,LLMs很少能达到人类顶级表现,并且在不同竞赛中表现稳定性较差。没有单一模型占据主导地位:少数竞赛由LLMs获胜,但人类在大多数任务上仍保持最高分。我们将GNN-CB作为一个动态基准发布,配备自动化评估基础设施、动态排行榜和可复现的执行流程。除基准测试外,GNN-CB还提供了一个面向实践的资源,用于研究跨渐进多样图学习任务的GNN实现。该基准和评估框架可在以下https URL公开获取。
英文摘要
Large language models (LLMs) have demonstrated strong performance on coding and reasoning benchmarks; however, their ability to solve graph-structured machine learning problems remains largely unexplored. In particular, no benchmark currently evaluates whether LLMs can autonomously solve end-to-end Graph Neural Network (GNN) coding tasks under realistic competition settings. To address this gap, this paper introduces GNN-CB, the first competition-based benchmark for evaluating both humans and LLMs on GNN coding tasks. GNN-CB consists of 18 curated competitions spanning node-, edge-, and graph-level prediction across diverse graph categories, domains, and difficulty tiers. All submissions are evaluated through a unified automated pipeline with hidden test sets and standardized scoring. Human participants solve tasks under controlled competition constraints, while LLMs are evaluated using a frozen zero-shot prompting protocol based on a plan-then-code paradigm with bounded execute-and-repair loops. The benchmark additionally supports both non-agent and autonomous agent-based evaluation within the same protocol. Under our evaluated protocol, LLMs rarely match Human Top performance and show less stable performance across competitions. No single model dominates: a few competitions are won by LLMs, yet humans still hold the top score on most tasks. We release GNN-CB as a living benchmark with automated evaluation infrastructure, dynamic leaderboards, and reproducible execution pipelines. Beyond benchmarking, GNN-CB provides a practice-oriented resource for studying GNN implementation across progressively diverse graph-learning tasks. The benchmark and evaluation framework are publicly available at https://basiralab.github.io/GNN-CB/.
CommentsAccepted at EMNLP 2026. 28 pages, 16 figures, 4 tables. Benchmark and leaderboards: https://basiralab.github.io/GNN-CB/