arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03816cs.SIcs.CVcs.LG

当视觉遇上图:图推理与学习综述

When Vision Meets Graphs: A Survey on Graph Reasoning and Learning

Xinjian Zhao, Wei Pang, Zhixuan Yu, Xiangru Jian, Xiaozhuang Song, Yaoyao Xu, Zhongkai Xue, Dingshuo Chen, Shu Wu, Philip Torr, Tianshu Yu

首次发表
浏览论文内容

中文总结 AI 辅助

本综述首次系统概述“视觉遇上图”领域,将图的视觉表征作为推理与学习的一级输入,梳理相关研究方向,旨在推进能像科学家一样处理图的基础模型。

中文摘要 AI 辅助

图是自然科学与社会科学诸多问题的基础数据结构。过去十年,图神经网络(GNN)凭借扎实的理论基础主导了图机器学习领域。然而,科学家常通过视觉理解图结构:化学家研读分子图,社会学家分析网络可视化。尽管图可视化研究已有数十年,多数图学习流程仍仅将图视为符号结构,极少利用图的视觉形式。我们认为,在视觉与视觉语言模型强大的当下,这一差距值得重新关注。本综述首次系统概述了新兴的“视觉遇上图”领域,该领域将图的视觉表征视为推理与学习的一级输入。我们将现有工作分为三个方向:图推理的视觉方法研究模型如何利用图的视觉表征理解结构并开展多步推理;图学习的视觉方法探索视觉特征如何在消息传递的已知局限之外补充或增强图编码器;科学图研究那些标准化表征规范同时支持推理与学习的领域。我们的目标是阐明当前方法能做与不能做的事,并规划一条路径,以构建能像科学家一样感知和推理图的基础模型。

英文摘要

Graphs are a fundamental data structure underlying many problems in the natural and social sciences. Over the past decade, Graph Neural Networks (GNNs) have dominated graph machine learning, supported by solid theoretical foundations. Yet scientists often understand graph structure through vision: chemists read molecular diagrams and social scientists inspect network visualizations. Despite decades of work on graph visualization, most graph learning pipelines still treat graphs purely as symbolic structures, rarely leveraging the visual form of graphs. We argue that this gap deserves renewed attention in the era of powerful vision and vision-language models. This survey provides a first systematic overview of the emerging area we term vision meets graphs, which treats visual depictions of graphs as first-class inputs for reasoning and learning. We organize existing work into three threads. Vision for Graph Reasoning studies how models can use visual depictions of graphs to understand structure and carry out multi-step reasoning. Vision for Graph Learning explores how visual features can complement or augment graph encoders beyond known limitations of message passing. Scientific Graphs examines domains where standardized depiction conventions support both reasoning and learning. Our goal is to clarify what current methods can and cannot do, and to outline a path toward foundation models that perceive and reason about graphs as scientists do.

发表机构

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • University of Waterloo(滑铁卢大学)
  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑