arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GNN-CB:面向人类与LLM评估的图神经网络竞赛基准

GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation

Murad Hossen, Tasneem Selim, Gurur Gamgam, Tuga Yousif, Abderrahmane Kasmi, Ikram Aissiou, Mubaraq Onipede, Faran Taimoor Butt, Sanae Zrigui, Rosa Y. G. Paccotacya-Yanque, Ignatius Balayo, Ikram Elhouiti, Hadil Affes, Bijay Adhikari, Sargam Goyal, Muhammad Ibrahim Isah, Mohammad Idrees Bhat, Samuel Kangoni Matia, Peguy Kem-Meka Tiotsop Kadzue, Maha Trabelsi, Emmanuel Owusu, Vinit, Nour Majdoub, Tamiru Alemnew, Islem Rekik

arXiv 2610.05387首次发表:更新:

发表机构

University of Houston; Alexandria University; Bogazici University; Ankara Yıldırım Beyazıt University; ESI, Algiers; University of Algiers 1; York St John University London; Air University, Islamabad; Mohammed Premier University; Universidad Católica San Pablo; Busitema University; University of Laghouat; University of Tunis El Manar; Tribhuvan University; IIT Roorkee; Shobhit Institute of Engineering and Technology; MIT World Peace University; University of Kinshasa; University of Bertoua; University of the Witwatersrand(休斯顿大学; 亚历山大大学; 博阿齐奇大学; 安卡拉耶尔德勒姆贝亚泽特大学; 阿尔及尔ESI; 阿尔及尔第一大学; 约克圣约翰大学伦敦校区; 伊斯兰堡空军大学; 穆罕默德第一大学; 圣巴勃罗天主教大学; 布西特马大学; 拉格瓦特大学; 突尼斯埃尔马纳尔大学; 特里布万大学; 印度理工学院鲁尔基分校; 肖比特工程技术学院; MIT世界和平大学; 金沙萨大学; 贝尔图阿大学; 威特沃特斯兰德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLMs在端到端GNN编码任务上能力评估缺失的问题,提出首个竞赛基准GNN-CB,含18个竞赛,统一自动化评估,发现LLMs难及人类顶级表现。

AI 中文摘要

大型语言模型(LLMs)在编码和推理基准测试中展现出强大性能;然而,它们解决图结构机器学习问题的能力在很大程度上仍未得到探索。特别是,目前尚无基准测试评估LLMs能否在现实竞赛设置下自主解决端到端的图神经网络(GNN)编码任务。为弥补这一空白,本文介绍了GNN-CB,这是首个基于竞赛的基准,用于评估人类和LLMs在GNN编码任务上的表现。GNN-CB包含18个精选竞赛,涵盖节点级、边级和图级预测,涉及多种图类别、领域和难度层级。所有提交均通过统一的自动化流程进行评估,该流程包含隐藏测试集和标准化评分。人类参与者在受控竞赛约束下解决问题,而LLMs则通过基于计划-然后-编码范式的冻结零样本提示协议进行评估,并带有有界的执行-修复循环。该基准还支持在同一协议内进行非智能体和基于自主智能体的评估。在我们评估的协议下,LLMs很少能达到人类顶级表现,并且在不同竞赛中表现稳定性较差。没有单一模型占据主导地位:少数竞赛由LLMs获胜,但人类在大多数任务上仍保持最高分。我们将GNN-CB作为一个动态基准发布,配备自动化评估基础设施、动态排行榜和可复现的执行流程。除基准测试外,GNN-CB还提供了一个面向实践的资源,用于研究跨渐进多样图学习任务的GNN实现。该基准和评估框架可在以下https URL公开获取。

英文摘要

Large language models (LLMs) have demonstrated strong performance on coding and reasoning benchmarks; however, their ability to solve graph-structured machine learning problems remains largely unexplored. In particular, no benchmark currently evaluates whether LLMs can autonomously solve end-to-end Graph Neural Network (GNN) coding tasks under realistic competition settings. To address this gap, this paper introduces GNN-CB, the first competition-based benchmark for evaluating both humans and LLMs on GNN coding tasks. GNN-CB consists of 18 curated competitions spanning node-, edge-, and graph-level prediction across diverse graph categories, domains, and difficulty tiers. All submissions are evaluated through a unified automated pipeline with hidden test sets and standardized scoring. Human participants solve tasks under controlled competition constraints, while LLMs are evaluated using a frozen zero-shot prompting protocol based on a plan-then-code paradigm with bounded execute-and-repair loops. The benchmark additionally supports both non-agent and autonomous agent-based evaluation within the same protocol. Under our evaluated protocol, LLMs rarely match Human Top performance and show less stable performance across competitions. No single model dominates: a few competitions are won by LLMs, yet humans still hold the top score on most tasks. We release GNN-CB as a living benchmark with automated evaluation infrastructure, dynamic leaderboards, and reproducible execution pipelines. Beyond benchmarking, GNN-CB provides a practice-oriented resource for studying GNN implementation across progressively diverse graph-learning tasks. The benchmark and evaluation framework are publicly available at https://basiralab.github.io/GNN-CB/.

CommentsAccepted at EMNLP 2026. 28 pages, 16 figures, 4 tables. Benchmark and leaderboards: https://basiralab.github.io/GNN-CB/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑