arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高效统一多模态理解(EUMU):第八届LSVOS挑战赛MUMU赛道获胜方案

Efficient Unified Multimodal Understanding (EUMU): Winning Solution for the MUMU Track at the 8th LSVOS Challenge

Dayoung Kil, Seong-heum Kim

arXiv 2609.19451首次发表:更新:

发表机构

Soongsil University(崇实大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EUMU是第八届LSVOS挑战赛MUMU赛道的获胜方案,通过共享预训练模型和跨任务线索细化,在单个模型中统一实现图像打标、开放词汇检测和图像描述,以239.169M参数和23.947 GFLOPs达到17.3409分。

AI 中文摘要

移动统一多模态理解(MUMU)挑战赛要求使用单个高效模型同时执行多概念图像打标、开放词汇目标检测和图像描述生成。我们提出了高效统一多模态理解(EUMU),这是第八届LSVOS挑战赛MUMU赛道的获胜方案。EUMU基于共享的预训练多模态模型,利用其基于提示的能力进行检测和描述,并在共享视觉特征上训练轻量级头部以预测质量、场景和事件标签。EUMU并未将三个任务独立处理,而是通过重用任务输出作为跨任务线索来应用任务感知推理细化。对于检测,描述线索有助于恢复初始检测遗漏的目标。对于描述生成,检测线索有助于细化描述以更好地反映检测到的目标。对于打标,图像统计信息细化质量预测,而描述和检测线索细化场景和事件预测。这种设计在满足挑战资源约束的同时,将三个任务统一到单个模型中。EUMU包含239.169M参数,需要23.947 GFLOPs,峰值推理内存为4.5 GB,最终挑战得分为17.3409。代码和模型可在该https URL获取。

英文摘要

The Mobile Unified Multimodal Understanding (MUMU) Challenge requires a single efficient model to jointly perform multi-concept image tagging, open-vocabulary object detection, and image captioning. We present Efficient Unified Multimodal Understanding (EUMU), the winning solution for the MUMU Track of the 8th LSVOS Challenge. EUMU builds on a shared pretrained multimodal model, using its prompt-based capabilities for detection and captioning and training lightweight heads on shared visual features to predict quality, scene, and event tags. Rather than treating the three tasks independently, EUMU applies task-aware inference refinement by reusing task outputs as cross-task cues. For detection, caption cues help recover objects missed by the initial detection. For captioning, detection cues help refine the caption to better reflect the detected objects. For tagging, image statistics refine quality predictions, while caption and detection cues refine scene and event predictions. This design unifies all three tasks within a single model while satisfying the challenge's resource constraints. EUMU contains 239.169M parameters, requires 23.947 GFLOPs, uses 4.5 GB of peak inference memory, and achieves a final challenge score of 17.3409. Code and models are available at https://github.com/Dayoung-Kil/EUMU.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑