AI 中文总结
该研究针对多模态大语言模型推理效率不足的问题,提出MMDynOpt-Agent,通过强化学习将多模态推理动态优化建模为马尔可夫决策过程,结合多轮优化提示与多维度奖励机制,在15个公开数据集上性能优于基线方法。
AI 中文摘要
近年来,多模态大语言模型(MLLM)在视觉理解和复杂推理任务中展现出强大潜力。然而,现有方法往往难以将多模态输入中的视觉线索与问题语义高效转化为有效的推理条件,从而限制了多模态大语言模型的推理性能。为应对这一挑战,我们提出MMDynOpt-Agent,该方法通过端到端强化学习将多模态推理的动态优化建模为马尔可夫决策过程。具体而言,一个轻量级多模态智能体作为决策策略,与作为环境的目标MLLM交互,通过多轮动态优化提示自适应引导其推理过程。此外,为降低多模态推理的成本,我们设计了一种结合格式合规性、答案正确性和预算感知的奖励机制,以共同确保推理的准确性和效率。MMDynOpt-Agent具有可迁移性和通用性,可在一个目标MLLM上进行训练,并在推理时迁移至其他MLLM。在15个公开数据集上的实验结果表明,MMDynOpt-Agent取得了优异的性能,且优于基线方法。我们的项目可通过此https URL获取。
英文摘要
Recently, multimodal large language models (MLLMs) have demonstrated strong potential in visual understanding and complex reasoning tasks. However, existing methods often struggle to efficiently transform visual cues from multimodal inputs and the semantics of the question into effective reasoning conditions, thereby limiting the reasoning performance of multimodal large language models. To address this challenge, we propose MMDynOpt-Agent, which models the dynamic optimization of multimodal reasoning as a Markov decision process via end-to-end reinforcement learning. Specifically, a lightweight multimodal agent serves as the decision policy and interacts with the target MLLM as the environment, adaptively steering its reasoning through multi-turn dynamic optimization prompts. Furthermore, to reduce the cost of multimodal reasoning, a reward mechanism that combines format compliance, answer correctness, and budget awareness is designed to jointly ensure reasoning accuracy and efficiency. MMDynOpt-Agent is transferable and generalizable, enabling training with one target MLLM and inference-time transfer to others. Experimental results on fifteen public datasets show MMDynOpt-Agent achieves strong performance and outperforms baselines. Our project is available at https://github.com/QwenQKing/MMDynOpt-Agent.