Rethinking Video-Language Model from the Language Input Perspective
从语言输入角度重新思考视频-语言模型
机构 * School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) ; Nanyang Technological University, Singapore(新加坡南洋理工大学) ; University College London(伦敦大学学院) ; Huazhong University of Science and Technology(华中科技大学) ; Wuhan University(武汉大学)
AI总结 本文从语言输入角度出发,提出一种即插即用的框架,通过生成正负文本、属性文本推理和自加权损失,提升视频-语言模型的性能。
Comments Published in AAAI 2026