发表机构
Peking University; Kling Team; HKUST(GZ); CUHK; ZJU; THU(北京大学; Kling团队; 香港科技大学(广州); 香港中文大学; 浙江大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对现有智能体视觉推理模型模式适应性有限、工具增益被损害抵消的问题,提出Beacon模型,通过强化学习相关机制提升性能与适应性,在多基准上表现优异。
AI 中文摘要
智能体视觉推理的核心目标是提升多模态大语言模型(MLLM)在复杂任务上的成功率,而非仅为其配备一套复杂却低效的推理范式。本研究从工具使用的两个关键维度重新审视智能体视觉推理:模式适应性(MA)与工具效果(TE)。模式适应性指MLLM能否识别工具的真实必要性并相应调用,从而避免不必要的计算开销,同时提升需要工具辅助的难题性能;工具效果指工具使用的实际影响:工具应扩展模型在仅靠文本推理无法解决的问题上的能力,同时避免在模型无需工具即可解决的问题上引入额外错误。我们开展综合分析以量化这两个属性,实证发现现有智能体视觉推理模型的模式适应性有限,且工具在难题上带来的增益很大程度上被在模型本可解决的简单问题上造成的损害所抵消。基于这些观察,我们提出Beacon,一种新型智能体视觉推理模型,可实现更强的整体性能、改进的模式适应性及真正的工具诱导性能增益。Beacon的核心是强化学习阶段的必要性感知自适应奖励和提示引导的能力扩展机制,二者分别鼓励基于任务必要性的自适应工具调用,并增强模型在最具挑战性问题上的工具使用能力。在不同基准上的大量实验表明,Beacon具有强劲的整体性能,且在模式适应性和工具效果两方面均有显著提升。
英文摘要
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks. We rethink agentic visual reasoning through two key dimensions of tool use: Mode Adaptiveness and Tool Effect. Mode Adaptiveness characterizes whether an MLLM recognizes when tools are necessary and invokes them accordingly, avoiding unnecessary computational overhead while improving performance on problems requiring tool assistance. Tool Effect characterizes whether tools extend the model's capabilities on problems unsolvable through tool-free reasoning without introducing errors on problems it can already solve. Our analysis quantifies these properties and reveals that existing models exhibit limited Mode Adaptiveness, while tool-use gains on hard examples are largely offset by harm on easy ones. Motivated by these observations, we propose Beacon, a novel agentic visual reasoning model trained with supervised fine-tuning (SFT) and reinforcement learning (RL). Its RL stage combines Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion. Necessity-Aware Adaptive Reward encourages tool-free solutions when they succeed while preserving full reward for successful tool use when tool-free rollouts fail. Hint-Guided Capability Expansion uses verified, answer-free expert hints to recover learning signals from all-wrong rollout groups, aiming to extend tool-use capability on the hardest problems. Across 13 benchmarks, Beacon achieves the highest average score among the evaluated open-source models and ranks first on 11 benchmarks. On five diagnostic benchmarks, it improves the average tool-available accuracy over its tool-free accuracy by 1.96 points and achieves the largest tool-gain minus tool-harm score (+3.14 points). These results show Beacon's advanced performance, Mode Adaptiveness, and the net benefit of tool use.
Comments35 pages