发表机构
University of Electronic Science and Technology of China; Tsinghua Shenzhen International Graduate School(电子科技大学; 清华大学深圳国际研究生院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出VGCompiler,通过表示编译器和操作编译器从图图像中恢复显式图表示并编译查询意图,利用强化学习训练,在多个基准上显著超越现有方法,并泛化至真实世界领域。
AI 中文摘要
视觉图推理要求直接从图图像中回答图论问题,其中图拓扑和状态通过视觉方式传达,而非以符号形式给出。尽管视觉语言模型(VLM)近期取得了进展,但当前的视觉图推理方法在简单的视觉图问题上仍然失败。这揭示了现有方法的一个根本性局限:它们优先考虑最终答案的监督,而忽视了从视觉输入中恢复显式图表示(该表示保留图拓扑和状态)的中间过程。为解决这一局限,我们提出VGCompiler,一种基于知识编译的视觉图推理编译中心范式。VGCompiler围绕两个编译器组织推理:一个表示编译器,从视觉输入中恢复保持结构的中间图表示;一个操作编译器,将恢复的图状态下的查询意图编译为可执行的图操作。具体而言,我们基于Qwen3-VL-8B构建VGCompiler,并使用由可执行性、编译图有效性、表示质量和操作质量的分层奖励引导的强化学习进行训练。VGCompiler使用冻结的观察器将图和问题条件总结为轻量级签名,从而支持跨相似场景的档案检索和代码复用。在GVLQA、VisionGraph和VGCURE三个基准上的实验表明,基于8B骨干构建的Qwen-VGCompiler比最强的闭源VLM基线高出28.9%,比最强的基于代码的基线高出23.7%,同时保持高效率。我们进一步在三个真实世界领域评估VGCompiler,包括地铁线路规划、物流配送和网络故障评估,它能够泛化到异构视觉图和领域相关任务。
英文摘要
Visual graph reasoning requires answering graph-theoretic questions directly from graph images, where graph topology and state are conveyed visually rather than given in symbolic form. Despite recent progress of vision-language models (VLMs), current approaches to visual graph reasoning still fail on simple visual graph problems. This reveals a fundamental limitation of existing approaches: they prioritize final-answer supervision over the intermediate recovery of an explicit graph representation that preserves graph topology and state from visual input. To address this limitation, we propose VGCompiler, a compilation-centric paradigm for visual graph reasoning via knowledge compilation. VGCompiler organizes reasoning around two compilers: a representation compiler that recovers a structure-preserving intermediate graph representation from visual input, and an operation compiler that compiles query intent under the recovered graph state into an executable graph operation. Specifically, we build VGCompiler on Qwen3-VL-8B and train it with reinforcement learning guided by a layered reward over executability, compiled graph validity, representation quality, and operation quality. VGCompiler uses a frozen observer to summarize graph and question conditions into lightweight signatures, enabling archive retrieval and code reuse across similar regimes. Experiments on three benchmarks GVLQA, VisionGraph, and VGCURE, show that Qwen-VGCompiler, built on an 8B backbone, surpasses the strongest closed-source VLM baseline by 28.9% and the strongest code-based baseline by 23.7%, while maintaining high efficiency. We further evaluate VGCompiler on three real-world domains, including metro routing, logistics delivery, and network fault assessment, where it generalizes across heterogeneous visual graphs and domain-grounded tasks.
CommentsAccepted at the 34th ACM International Conference on Multimedia (MM '26)