First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
机构 * School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院) ; Zhongguancun Academy(中关村学院) ; Shanghai Innovation Institute(上海创新研究院) ; Lehigh University(莱特大学)
专题命中 多模态生成 :multi-modal(title,abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted by NeurIPS 2025