Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos
Sa2VA:将SAM2与多模态大语言模型结合用于图像与视频的密集接地理解
机构 * University of California, Merced(加州大学默塞德分校) ; Bytedance Seed(字节跳动种子) ; Wuhan University(武汉大学) ; Peking University(北京大学)
专题命中 多模态训练与对齐 :MLLM(title,summary_cn);multi-modal(abstract);分类 cs.CV
AI总结 本研究提出Sa2VA,结合SAM2与MLLM实现图像视频密集接地理解,引入Ref-SAV数据集,在多任务中表现优异且可扩展至多款开源MLLM,代码模型已公开。
Comments Accepted by IEEE TPAMI. Code: this https URL (https://github.com/Bytedance/Sa2VA)