arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06399cs.AI

CVIF:面向多模态大语言模型几何图解理解的关键性驱动视觉干预框架

CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs

  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

Jiahui Kang, Bifan Wei, Lingling Zhang, Tianwen Jiang, Qiuyong Xiao, Jihong Zhang, Jun Liu

AI总结:

针对多模态大语言模型在几何图解理解中因稀疏视觉线索和模糊关联而依赖文本先验的问题,提出无需训练的关键性驱动视觉干预框架(CVIF),通过定位关键层并执行视觉干预,在PGPS9K和PGDP5K上显著提升总体F1分数,建立了推理时视觉干预新范式。

AI中文摘要:

尽管多模态大语言模型(MLLMs)在视觉任务上取得了显著进展,但由于稀疏视觉线索和模糊的符号-基元关联的存在,几何图解理解仍然具有挑战性。因此,MLLMs可能依赖文本先验,产生与视觉证据相冲突的解释。我们引入了无需训练的“关键性驱动视觉干预框架”(CVIF),这是一种推理时方法,在从证据聚合到语义解码的过渡期间定位关键层并执行视觉干预。在这些层中,一个“几何约束局部关系重建”(GCLR)模块选择并加权以顶点为中心的视觉证据,而一个“自适应视觉转向算子”(AVSO)将注意力质量重新分配给选定的标记。在PGPS9K和PGDP5K上的实验表明,CVIF将总体F1分数分别从77.85提高到85.58,从75.23提高到82.84,建立了一种新颖的推理时视觉干预范式。

英文摘要:

Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse visual cues and ambiguous symbol-primitive associations. MLLMs may therefore rely on textual priors, producing interpretations that conflict with visual evidence. We introduce the training-free Criticality-Driven Visual Intervention Framework (CVIF), an inference-time method that localizes critical layers and executes visual interventions during the transition from evidence aggregation to semantic decoding. At these layers, a Geometry-Constrained Local Relation Reconstruction (GCLR) module selects and weights vertex-centered visual evidence, while an Adaptive Visual Steering Operator (AVSO) redistributes attention mass toward the selected tokens. Experiments on PGPS9K and PGDP5K show that CVIF raises Overall F1 from 77.85 to 85.58 and from 75.23 to 82.84, respectively, establishing a novel inference-time visual intervention paradigm.

↑