Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
通过边界框思考:通过强化微调提升空间时间视频定位
机构 * ByteDance Intelligent Creation(字节跳动智能创作) ; Tsinghua University(清华大学) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; Shanghai Jiao Tong University(上海交通大学) ; Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
AI总结 STVG-o1通过强化微调提升空间时间视频定位性能,实现无需架构修改的先进结果。