大型视觉语言模型中的认知专家混合体
Mixture of Cognitive Experts in Large Vision-Language Models
浏览论文内容
中文总结 AI 辅助
研究大型视觉语言模型中多编码器专家整合难题,提出证据驱动的多模态推理框架,利用布鲁姆分类法及两阶段认知语言化,改进感知和推理能力,实现细粒度分析。
中文摘要 AI 辅助
大型视觉语言模型(LVLMs)需要对视觉和文本输入进行强大的推理。近期研究表明认知元素与更好的性能相关。许多所需的感知功能已由特定领域的计算机视觉模型提供,关键挑战是将这些多编码器专家整合为可信赖、可解释且连贯的表示。我们提出了一个证据驱动的多模态推理框架,利用受布鲁姆启发的分类法作为分层推理协议。两阶段认知语言化首先通过将专家输出分解为简短的原子证据陈述来生成文字证据摘要,然后进行布鲁姆语言化将这些证据项转化为分阶段的推理轨迹,一个轻量级推理轨迹模块对轨迹进行定量分析以使证据使用和推理过程明确。通过这种整合,在感知和推理能力方面有了一些改进,并且轨迹模块提供了定量证据表明不同查询会引发不同的认知进入水平和证据使用轨迹,从而实现细粒度分析。
英文摘要
Large Vision Language Models (LVLMs) require strong reasoning over both visual and textual input. Recent work suggests that cognitive elements, especially diverse representations and metacognition, correlate with better performance. Many of the needed perceptual functions are already provided by specialized domain-specific computer vision models, which act as the perceptual subsystem for detecting objects, localizing them, inferring states, recovering spatial layout, and reading text. The key challenge is to integrate these multi-encoder experts into a trustworthy, interpretable, and coherent representation that improves verifiability and reduces hallucinations. This is difficult because vision-language questions span different cognitive levels, yet most LVLM pipelines apply the same perception-reasoning routing regardless of the demand of each query. We propose an evidence-driven multimodal reasoning framework that utilizes a Bloom-inspired taxonomy as a hierarchical reasoning protocol. The two-stage cognitive verbalization first produces a Literal Evidence Summary by decomposing expert outputs into short, atomic evidence statements. It then performs Bloom Verbalization to turn these evidence items into a staged reasoning trace, and a lightweight Reasoning Trace Module quantitatively analyzes the trace to make evidence usage and reasoning progression explicit. Through this integration, we observed several improvements in perception and reasoning abilities. Moreover, the trace module provides quantitative evidence that different queries induce different cognitive entry levels and evidence-use trajectories that enable fine-grained analysis.
发表机构
- Singapore University of Technology and Design (SUTD)(新加坡科技设计大学)
机构由 AI 辅助整理,请以论文原文为准。