MoTE: Mixture of Task Experts for Multi-Task Video Understanding
MoTE:面向多任务视频理解的任务专家混合模型
机构 * University of Kaiserslautern-Landau (RPTU)(凯泽斯劳滕-兰道大学(RPTU)) ; German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
专题命中 视频理解 :video understanding(title);video-language(abstract);分类 cs.CV
AI总结 针对多任务视频理解中现有解码器的任务行为纠缠、能力扩展难等问题,提出MoTE架构,实例化为VideoLLM-MoTE,在COIN基准上表现优于基线,实现了可解释且计算高效的多任务视频-语言学习。
Comments Accepted at BMVC 2026. 32 pages, 4 figures, 15 tables, including supplementary material