Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions
面向具有属性结构和质量验证指令的通用视频MLLMs
机构 * VCIP, School of Computer Science, Nankai University(VCIP,计算机科学学院,南开大学) ; ByteDance Inc.(字节跳动公司) ; Tsinghua University(清华大学)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
AI总结 本研究提出ASID-1M数据集、ASID-Verify验证流程和ASID-Captioner模型,通过细粒度结构化指令提升视频理解性能,实现高质量描述生成与指令遵循。
Comments Project page: https://asid-caption.github.io/