MoTE:面向多任务视频理解的任务专家混合模型
MoTE: Mixture of Task Experts for Multi-Task Video Understanding
- University of Kaiserslautern-Landau (RPTU)(凯泽斯劳滕-兰道大学(RPTU))
- German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多任务视频理解中现有解码器的任务行为纠缠、能力扩展难等问题,提出MoTE架构,实例化为VideoLLM-MoTE,在COIN基准上表现优于基线,实现了可解释且计算高效的多任务视频-语言学习。
AI中文摘要:
过程视频-语言模型必须基于相同的视觉证据解决异构任务,包括动作识别、预测和过程预测。密集Transformer解码器在各任务间共享相同的前馈网络,这会使任务行为纠缠,难以实现可控的能力扩展。稀疏专家混合(MoE)解码器提供条件计算,但基于token的学习路由无法自然适配任务级过程目标。我们提出MoTE(任务专家混合模型),一种将大语言模型前馈网络转换为任务特定专家、同时保持多模态骨干共享的解码器架构。每个样本遵循一条样本级任务路由,因此活跃的任务专家计算与存储的任务专家数量无关。我们将该设计实例化为VideoLLM-MoTE,在五个COIN基准上使用显式任务路由进行评估。拥有五个专家的模型每个样本激活约20亿个LLM参数,比近期VideoLLM基线实现更高的平均top-1准确率。在相同专家拓扑下,它优于密集全专家激活和学习到的稀疏路由控制。这些结果表明,任务结构化路由为多任务视频-语言学习提供了一种可解释且计算高效的解码器替代方案。
英文摘要:
Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedure prediction. Dense transformer decoders share the same feed-forward networks across tasks, which can entangle task behavior and make controlled capability expansion difficult. Sparse Mixture-of-Experts (MoE) decoders provide conditional computation, but token-level learned routing is not naturally aligned with task-level procedural objectives. We propose MoTE (Mixture of Task Experts), a decoder architecture that converts large language model feed-forward networks into task-specific experts while keeping the multimodal backbone shared. Each example follows one sample-level task route, so active task-expert computation remains independent of the number of stored task experts. We instantiate this design as VideoLLM-MoTE and evaluate it on five COIN benchmarks using explicit task routes. The five-expert model activates ~2B LLM parameters per sample and achieves higher average top-1 accuracy than recent VideoLLM baselines. Under the same expert topology, it improves over dense all-expert activation and learned sparse-routing controls. These results show that task-structured routing provides an interpretable and compute-efficient decoder alternative for multi-task video-language learning.