CURV:通过课程可视化接地推理增强图表理解
CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
浏览论文内容
中文总结 AI 辅助
针对多模态大语言模型视觉接地与推理不足的问题,提出CURV课程学习框架,结合CCQA数据集,在图表问答任务中实现显著性能提升并具备良好泛化性。
中文摘要 AI 辅助
图表问答(CQA)要求多模态大语言模型(MLLMs)整合视觉理解与逻辑推理,但现有模型在精准视觉接地和连贯推理链方面存在不足。尽管外在思维链提示和视觉线索可显著提升性能,当前MLLMs仍缺乏内在的视觉接地推理能力,导致感知不准确且推理与视觉证据脱节。为解决这些局限,我们提出CURV,这是一个课程学习框架,通过将CQA重新表述为多步骤视觉接地推理来开发内在视觉推理能力,其中每一步通过空间注意力集中协调逻辑推理与动态视觉接地。为辅助模型学习,我们进一步引入CCQA,这是一个三级课程数据集,可针对不同图表类型和推理模式进行可扩展的合成生成。我们的课程从基础的单步推理系统推进到复杂的多图表组合任务。实验表明,CURV相较于基准模型可实现最高20.50%的提升,且可泛化到真实世界基准(最高提升12.30%)和域外多模态推理任务(最高提升10.20%),验证了将动态接地的视觉推理内在化以增强图表理解能力的有效性。代码可在该URL获取。
英文摘要
Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains. While extrinsic chain-of-thought prompting and visual cues significantly improve performance, current MLLMs lack intrinsic visual grounded reasoning capabilities, leading to inaccurate perception and reasoning disconnected from visual evidence. To address these limitations, we propose CURV, a curriculum learning framework that develops intrinsic visual reasoning capabilities by reformulating CQA as multi-step visual grounded reasoning, where each step coordinates logical reasoning with dynamic visual grounding through spatial attention concentration. To assist model learning, we further introduce CCQA, a three-level curriculum dataset with scalable synthetic generation across diverse chart types and reasoning patterns. Our curriculum systematically progresses from basic single-operation reasoning to complex multi-chart compositional tasks. Experiments demonstrate that CURV achieves up to $\uparrow20.50\%$ improvements over baselines and is generalizable to real-world benchmarks (up to $\uparrow12.30\%$) and out-of-domain multimodal reasoning tasks (up to $\uparrow10.20\%$), validating the effectiveness of internalizing visual reasoning with dynamic grounding for enhanced chart understanding capabilities. Code is available at: https://xhguo7.github.io/CURV/.