发表机构
Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大语言模型忽视结构化关系的问题,提出场景图思维(SaGe)范式,通过自动数据引擎转换数据、构建训练数据,采用两阶段图对齐训练范式,在多模态基准测试中取得显著改进,提升细粒度感知和推理能力。
AI 中文摘要
多模态大语言模型展现出强大感知和推理能力,但多数现有模型专注孤立对象,忽视结构化关系,限制其在视觉密集任务中的表现。为应对此挑战,我们引入场景图思维(SaGe)这一新范式,通过显式场景图表示实现细粒度和结构化视觉推理。具体而言,先引入自动数据引擎将平面图像文本语料库转换为结构化场景图,基于此构建高质量训练数据,再引入两阶段图对齐训练范式。通过精心策划的数据和图对齐训练,该方法在八个多模态基准测试中取得显著改进,在细粒度感知和推理任务中展现出强大效果。
英文摘要
Multimodal Large Language Models (MLLMs) have demonstrated strong perception and reasoning capabilities. However, most existing models focus on isolated objects and neglect structured relationships for efficient target navigation, limiting their performance on visually intensive tasks. To address this challenge, we introduce Scene Graph Thinking (SaGe), a novel paradigm that enables fine-grained and structured visual reasoning through explicit scene-graph representations. Specifically, we first introduce an automated data engine that converts flat image-text corpora into structured scene graphs, where hierarchical entities constitute the nodes and diverse visual relations define the edges. Building upon this, we construct 120K high-quality training data by sampling reasoning traces from scene graphs. Then, two-stage graph-aligned post-training paradigms are introduced, where supervised fine-tuning internalizes MLLMs with structured reasoning, and subsequent reinforcement fine-tuning proposes node-as-proxy graph rewards to consolidate efficient graph exploration. With curated data and graph-aligned training, our approach achieves significant improvements across eight multimodal benchmarks, demonstrating strong effectiveness on fine-grained perception and reasoning tasks. Code is available at https://github.com/zwyang6/SaGe.
CommentsICML 2026