超越视觉:通过自我调节的隐式视觉工具实现高效多模态推理
Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools
- Sun Yat-sen University(中山大学)
- The Hong Kong Polytechnic University(香港理工大学)
- Yinwang Intelligent Technology Co., Ltd.(银望智能科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对多模态大语言模型推理效率问题,提出超越视觉(BEE)的隐式视觉工具范式,通过将视觉工具调用行为纳入训练目标,经两阶段训练,包括形式化思维链监督微调与自我调节奖励驱动对齐,提升了模型在细粒度视觉感知任务中的性能和推理效率。
AI中文摘要:
近期多模态大语言模型在“图像思考”范式下的细粒度感知任务取得进展,但依赖频繁外部工具调用和重复图像重新编码,导致计算开销和推理延迟大。为此提出超越视觉(BEE)的隐式视觉工具范式,将视觉工具调用行为纳入训练目标,鼓励模型发展自我调节调用机制。其分两阶段训练:形式化思维链监督微调激活隐式工具表征和自适应切换能力;自我调节奖励驱动对齐,引入净工具增益量化冗余工具使用现象并提出奖励机制惩罚无效工具依赖。BEE在细粒度视觉感知中达最优性能,在一般推理任务中具竞争力且推理效率大幅提升。
英文摘要:
Recent multimodal large language models (MLLMs) have made remarkable progress on fine-grained perception tasks under the "Thinking with Images" (TwI) paradigm by iteratively performing various visual tool operations. However, this paradigm relies heavily on frequent external tool calls and repeated image re-encoding, which leads to substantial computational overhead and inference latency. To address these issues, we propose Beyond the Eye (BEE), a novel implicit visual tool paradigm centered on self-regulated capability. BEE directly incorporates visual tool invocation behaviors into the training objective and encourages the model to develop a self-regulated invocation mechanism. This design enables the model to adaptively balance internal knowledge and implicit tools, avoiding redundant tool usage while substantially reducing inference latency. Specifically, BEE involves a two-stage training process: (1) Formalized Chain-of-Thought (CoT) Supervised Fine-tuning (SFT). We construct CoT trajectories with structured tool slots and mixed invocation states. This stage activates the model's implicit tool representations and adaptive switching capability. (2) Self-regulated Reward-Driven Alignment. To address redundant tool usage caused by ambiguous cognitive boundaries, we first introduce the Net Tool Gain (NTG) metric to quantify this phenomenon. Based on this observation, we further propose a self-regulated reward mechanism. This mechanism penalizes ineffective tool dependency and encourages the model to perform knowledge routing, ensuring that implicit tools are invoked only when the model's internal knowledge is insufficient. BEE achieves state-of-the-art performance in fine-grained visual perception while remaining competitive in general reasoning tasks and achieving substantial gains in inference efficiency.