强化微调赋能多模态大语言模型的推理能力
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
- Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文为立场论文,主张强化微调(RFT)能增强多模态大语言模型(MLLMs)的推理能力,归纳了五个关键改进点,并提出五个未来研究方向。
AI中文摘要:
站在2025年,正值追求通用人工智能(AGI)的关键节点,强化微调(RFT)已展现出显著提升大语言模型(LLMs)推理能力的潜力,并推动了如OpenAI-o1和DeepSeek-R1等前沿AI模型的发展。此外,如何高效应用RFT以增强多模态大语言模型(MLLMs)的推理能力,已引起学术界的广泛关注。在这篇立场论文中,我们主张强化微调能够赋能多模态大语言模型的推理能力。首先,我们详细介绍了对该领域感兴趣的研究者应熟悉的基础背景知识。其次,我们将RFT在增强MLLMs推理能力方面的改进细致地归纳为五个关键点:多样化的模态、多样化的任务与领域、更优的训练算法、丰富的基准测试以及蓬勃发展的工程框架。最后,我们提出了社区未来可能考虑的五个有前景的研究方向。我们希望这篇立场论文能在迈向AGI的关键阶段为社区提供有价值的见解。关于RFT用于MLLMs的工作总结可在https://github.com/Sun-Haoyuan23/Awesome-RL-based-Reasoning-MLLMs获取。
英文摘要:
Standing in 2025, at a critical juncture in the pursuit of Artificial General Intelligence (AGI), reinforcement fine-tuning (RFT) has demonstrated significant potential in enhancing the reasoning capability of large language models (LLMs) and has led to the development of cutting-edge AI models such as OpenAI-o1 and DeepSeek-R1. Moreover, the efficient application of RFT to enhance the reasoning capability of multimodal large language models (MLLMs) has attracted widespread attention from the community. In this position paper, we argue that reinforcement fine-tuning powers the reasoning capability of multimodal large language models. To begin with, we provide a detailed introduction to the fundamental background knowledge that researchers interested in this field should be familiar with. Furthermore, we meticulously summarize the improvements of RFT in powering reasoning capability of MLLMs into five key points: diverse modalities, diverse tasks and domains, better training algorithms, abundant benchmarks and thriving engineering frameworks. Finally, we propose five promising directions for future research that the community might consider. We hope that this position paper will provide valuable insights to the community at this pivotal stage in the advancement toward AGI. Summary of works done on RFT for MLLMs is available at https://github.com/Sun-Haoyuan23/Awesome-RL-based-Reasoning-MLLMs.