arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TopoAgent:基于大型视觉语言模型的感知到推理的结构感知框架,用于图到图拓扑提取

TopoAgent: A Structure-Aware Perception-to-Reasoning Framework for Diagram-to-Graph Topology Extraction with Large Vision-Language Models

Bangwei Guo, Xujiang Zhao, Yanchi Liu, Wei Cheng, Shengyu Chen, Dongyue Li, Masaharu Morimoto, Takayuki Kuroda, Dimitris Metaxas, Haifeng Chen

arXiv 2608.28701首次发表:更新:

发表机构

Rutgers University; Meta; NEC Labs America; University of Electro-Communications; NEC Corporation(罗格斯大学; Meta; 美国NEC实验室; 电气通信大学; 日本电气公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对图到图拓扑提取任务构建了TopoBench-180基准,并提出TopoAgent框架,结合感知、结构先验与推理,在基准上性能优于现有模型,填补了多模态结构化理解的空白。

AI 中文摘要

图到图拓扑提取旨在从结构示意图中提取实体及其连接关系构成的图,该任务对当前视觉语言模型仍具挑战性,因为它既需要细粒度的感知定位,又需要具备全局一致性的拓扑感知推理能力。我们提出了TopoBench-180,这是一个经人工验证的图到图拓扑提取基准,以及TopoAgent,这是一种利用大型视觉语言模型实现可靠拓扑提取的结构感知感知到推理框架。TopoBench-180包含180幅结构示意图,涵盖网页风格和网络风格两类,附带规范的图标注。TopoAgent通过结合定位感知、全局结构先验、规范节点清单构建、以节点为中心的从局部到全局的关系推理以及拓扑一致性执行,逐步提取目标图。在TopoBench-180上的实验表明,TopoAgent的性能优于强大的视觉语言模型基线和近期的视觉推理框架,尤其在边提取方面表现突出。更广泛地说,这项工作通过建立图到图拓扑提取的基准和框架,填补了多模态结构化理解领域的重要空白。该基准及相关资源将在此httpsURL上公开发布。

英文摘要

Diagram-to-graph topology extraction aims to extract a graph of entities and their connections from a structural diagram. This task remains challenging for current vision-language models because it requires both fine-grained perceptual grounding and topology-aware reasoning with global consistency. We present TopoBench-180, a human-verified benchmark for diagram-to-graph topology extraction, and TopoAgent, a structure-aware perception-to-reasoning framework for reliable topology extraction using large vision-language models. TopoBench-180 contains 180 structural diagrams spanning Web-style and Network-style categories, paired with canonical graph annotations. TopoAgent progressively extracts the target graph by combining grounded perception, global structural priors, canonical node inventory construction, node-centric local-to-global relation reasoning, and topological consistency enforcement. Experiments on TopoBench-180 show that TopoAgent outperforms strong vision-language model baselines and recent visual reasoning frameworks, especially on edge extraction. More broadly, this work fills an important gap in multimodal structured understanding by establishing a benchmark and framework for diagram-to-graph topology extraction. The benchmark and associated resources will be publicly released at https://huggingface.co/datasets/WayneGuo0011/TopoBench-180.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑