3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
3D-RFT:基于视频的3D场景理解的强化微调
机构 * University of Science and Technology of China(中国科学技术大学)
AI总结 提出3D-RFT框架,将可验证奖励的强化学习(RLVR)扩展到视频3D感知与推理,通过直接优化评估指标(如3D IoU和F1分数)提升性能,4B模型超越8B模型。
Comments Accepted at ICML 2026. Project page: https://3d-rft.github.io/