发表机构
Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉问答中文本转换效率低的问题,提出VisKG强化学习框架,将视觉内容转为问题特定知识图谱,过滤噪声并保留推理结构,采用GDPO优化,在多个基准上以更少令牌达到更优性能。
AI 中文摘要
近期在视觉问答领域的研究表明,通过将视觉输入转换为文本表示,视觉-语言模型能够展现出强大的推理能力。这种转换的有效性取决于视觉细节保留的程度;模型需要呈现并对齐足以支持推理的显性和隐性知识,同时避免引入虚假假设。现有利用详细图像描述的方法会引入与推理任务无关的视觉细节,导致输入令牌数量膨胀并增加计算成本。为解决这些挑战,我们提出了VisKG,一个强化学习(RL)框架,其中模型学习将视觉内容转换为问题特定的知识图谱(KG)表示。该过程在遵循最小充分信息原则的同时,过滤掉感知噪声,保留链式思维推理所需的实体-关系结构。为确保RL后训练的稳定性,VisKG采用了组奖励解耦归一化策略优化(GDPO)。此外,我们通过负推理样本加强了监督阶段,在RL后训练之前让模型接触错误的推理路径。在科学、数学和通用视觉理解基准上的实验结果表明,VisKG达到了与基线相当或更好的性能,同时所需令牌数少于基于描述的表示。此外,使用GDPO训练VisKG相比其GRPO训练的对应版本,平均准确率提高了2%。这些结果表明,知识图谱表示是支持多步推理的一种有前景的方法,并为未来针对给定任务自适应选择最合适表示的研究开辟了道路。
英文摘要
Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs into textual representations. The effectiveness of this translation depends on how well visual details are retained; models need to surface and align both explicit and implicit knowledge sufficient to support reasoning, without introducing spurious assumptions. Existing methods that leverage detailed image captions introduce visual details unrelated to the reasoning task, inflating input token counts and increasing computational cost. To address these challenges, we propose VisKG, a reinforcement learning (RL) framework in which models learn to translate visual content into question-specific knowledge graph (KG) representations. This process filters out perceptual noise while preserving the entity-relation structure needed for chain-of-thought reasoning, following the principle of minimum sufficient information. To ensure stable RL post-training, VisKG adopts Group reward-Decoupled Normalization Policy Optimization (GDPO). In addition, we strengthen the supervision stage with negative rationale samples, exposing the model to incorrect reasoning paths before RL post-training. Experimental results across science, mathematics, and general visual understanding benchmarks show that VisKG achieves performance comparable to or better than baselines, while requiring fewer tokens than caption-based representations. Moreover, training VisKG with GDPO improves accuracy by 2% over its GRPO-trained counterpart on average. These results suggest that KG representations are a promising approach for supporting multi-step reasoning and open up future work on adaptively selecting the most suitable representation for a given task.