基于一致性的视听分割:第八届LSVOS挑战赛MeViS音频赛道冠军报告
Agreement-Based Audio-Visual Segmentation:Champion Report for the MeViS-Audio Track in the 8th LSVOS Challenge
浏览论文内容
中文总结 AI 辅助
本文针对第八届LSVOS挑战赛MeViS音频赛道,提出分阶段视听分割方案,以Qwen3-ASR为基础,通过掩码一致性选择轨迹并修正查询,结合多模态得分判断目标存在性,最终取得赛道第一的成绩。
中文摘要 AI 辅助
MeViS音频赛道要求系统对视频中口语化动作描述所对应的物体进行分割,当描述的目标不存在时返回空掩码。本文提出一种简单的分阶段解决方案:Qwen3-ASR首先将语音转换为文本;随后通过互补的定位与分割模型生成多个视频掩码轨迹;为避免依赖单一预测,选择与其他候选轨迹平均掩码一致性最高的轨迹;针对需要多目标跟踪的查询,通过少量明确的方向、数量及复数规则进行修正;最后,视频级分类器结合视觉、视听及视频内查询得分,判断是否存在目标。提交的系统取得了0.5952的J&F指标、0.7931的无目标准确率、0.9205的目标准确率,最终得分为0.769589,挑战主办方告知该结果在赛道中排名第一。
英文摘要
The MeViS-Audio track asks a system to segment the objects described by a spoken motion expression throughout a video and to return empty masks when the described target is absent. We present a simple staged solution. Qwen3-ASR first converts speech into text. Several video mask tracks are then produced with complementary grounding and segmentation models. Instead of trusting a single prediction, we select the track that has the highest average mask agreement with the other candidates. A small set of explicit direction, count, and plural rules corrects queries that require more than ordinary single-object tracking. Finally, a video-level classifier combines visual, audio-visual, and within-video query scores to decide whether any target is present. The submitted system obtains 0.5952 J &F, 0.7931 no-target accuracy, 0.9205 target accuracy, and a final score of 0.769589. The challenge organizers notified our team that this result ranked first in the track.
发表机构
- National 111 Project Base of Intelligent Information Processing(智能信息处理国家111计划基地)
机构由 AI 辅助整理,请以论文原文为准。