CapMap-MS-TTA:ECCV 2026 第八届 LSVOS 挑战赛 MUMU 赛道第三名解决方案
CapMap-MS-TTA: 3rd Place Solution for the MUMU Track of the 8th LSVOS Challenge at ECCV 2026
查看机构详情
- Netease YiDun AI Lab(网易易盾AI实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出基于 Florence-2 的无训练方案 CapMap-MS-TTA,通过关键词映射和多尺度翻转 TTA 统一解决 MUMU 赛道三项任务,分数从 15.16 提升至 16.4815,获第三名。
中文摘要 AI 辅助
第八届大规模视频对象分割(LSVOS)挑战赛的 MUMU 赛道要求一个统一的多模态模型在严格资源约束(参数不超过 0.5B,峰值 GPU 内存不超过 8 GB)下同时解决图像打标签(任务 A)、开放词汇目标检测(任务 B)和英文描述生成(任务 C)。我们提出 CapMap-MS-TTA,一种基于微软 Florence-2-base(约 231M 参数)的无训练提交方案,结合了描述关键词映射与多尺度翻转测试时增强。任务 C 使用原生 <DETAILED_CAPTION> 路径,并进行长度/标记清理。任务 A 通过扩展的关键词词典(采用整词匹配)和轻量级扩展提示阶段,将相同的详细描述映射到官方的质量/场景/事件词汇表。任务 B 运行 Florence-2 开放检测(<OD>),采用多尺度和水平翻转测试时增强(TTA),随后进行标签感知的非极大值抑制(NMS)。无需微调,该系统将我们复现的 Florence-2 基线从 15.16 提升至最佳公开分数 16.4815,并在最终 MUMU 排行榜上位列第三。
英文摘要
The MUMU track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge requires a single unified multimodal model to jointly solve image tagging (Task A), open-vocabulary object detection (Task B), and English captioning (Task C) under strict resource constraints (<=0.5B parameters and <=8 GB peak GPU memory). We present CapMap-MS-TTA, a training-free submission built on Microsoft Florence-2-base (~231M parameters), combining caption keyword mapping with multi-scale flip test-time augmentation. Task C uses the native <DETAILED_CAPTION> pathway with length/token sanitization. Task A maps the same detailed caption into the official quality/scene/event vocabularies via an expanded keyword lexicon with whole-word matching and a lightweight expand-hints stage. Task B runs Florence-2 open detection (<OD>) with multi-scale and horizontal-flip test-time augmentation (TTA), followed by label-aware non-maximum suppression (NMS). Without fine-tuning, the system improves our reproduced Florence-2 baseline from 15.16 to a best public score of 16.4815, and ranks 3rd on the final MUMU leaderboard.