arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从语音到掩码轨迹的运动感知推理:2026年第8届LSVOS挑战赛MeViS音频赛道亚军方案

Motion-Aware Reasoning from Speech to Mask Tracks: Runner-up Solution for the MeViS-Audio Track of the 8th LSVOS Challenge 2026

Jinxing Zhou, Suiyi Zhao, Yanghao Zhou, Ruohao Guo

arXiv 2608.22337首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence; Anhui University of Science and Technology; National University of Singapore; China Agricultural University(穆罕默德·本·扎耶德人工智能大学; 安徽理工大学; 新加坡国立大学; 中国农业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Speech2MaskTrack方案,整合语音识别、运动定位等技术,在第8届LSVOS挑战赛MeViS音频赛道获亚军,实现语音引导的参考视频对象分割任务。

AI 中文摘要

语音引导的参考视频对象分割旨在恢复由口语运动描述指定的对象的掩码轨迹。在此场景中,语音承载的是语言指令而非发声对象的声学证据,因此解决方案需整合语音识别、以运动为中心的时间定位、掩码跟踪以及明确的无目标处理。我们提出Speech2MaskTrack,作为第8届LSVOS挑战赛MeViS音频赛道的方案。Speech2MaskTrack将口语查询转录并编译为类别、数量、方向、交互角色和时间阶段的结构化约束。SAM3.1枚举多个实例轨迹,TRACE利用完整轨迹的运动和关系证据对其进行排序。冻结的词汇存在门可抑制排序后的SAM3.1基础预测;当该门预测目标存在时,可用的全表达式条件SaSaSa2VA轨迹会替换SAM3.1的掩码。仅保持为空的输出会进入GPT辅助恢复阶段,该阶段会在查询和掩码级验证下再次调用SaSaSa2VA。Speech2MaskTrack在官方挑战赛排名中获得第二名。

英文摘要

Speech-guided referring video object segmentation aims to recover the mask tracks of objects specified by a spoken motion description. Here, speech carries a linguistic instruction rather than acoustic evidence from a sounding object, so a solution must connect speech recognition, motion-centric temporal grounding, mask tracking, and explicit no-target handling. We introduce Speech2MaskTrack, our approach for the MeViS-Audio track of the 8th LSVOS Challenge. Speech2MaskTrack transcribes the spoken query and compiles it into structured constraints over category, count, direction, interaction role, and temporal phase. SAM3.1 enumerates multiple instance tracks, which TRACE ranks using complete-trajectory motion and relation evidence. A frozen lexical presence gate may suppress the ranked SAM3.1 base prediction. When the gate predicts that a target is present, an available full-expression-conditioned SaSaSa2VA track replaces the SAM3.1 mask. Only outputs that remain empty enter GPT-assisted recovery, which invokes SaSaSa2VA again under query- and mask-level verification. Speech2MaskTrack achieved second place in the official challenge ranking.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑