arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03964eess.AS

基于QIANGDA与VOXBLINK2-AVSE的身份保真视听目标说话人提取

Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE

Peijun Yang, Zhan Jin, Xiaoyi Qin, Ruiyi Gan, Hao Wang, Juan Liu, Ming Li

中文总结 AI 辅助

本文提出QIANGDA普通话AV-TSE基准与VOXBLINK2-AVSE数据集,采用含AV-HuBERT等的提取器,通过双指标评估,取得了视听目标说话人提取的良好性能。

中文摘要 AI 辅助

视听目标说话人提取应返回视频指示的说话人,但分离器可能忽略视觉线索,反复输出声学主导的语音。本文引入QIANGDA,这是一个普通话AV-TSE基准,包含联合录制的真实双人混合语音及同步多视角视频,每个场景还包含前置的仅A和仅B阶段,提供场景内说话人参考,共77个场景、7598个片段(11.84小时),含6042个双标注混合语音,处理后剩余6038个可评估混合语音和12076个目标说话人行。本文还从VoxBlink2整理出VOXBLINK2-AVSE,包含28421个身份的250828对同步音频-唇ROI,共766.17小时语音。本文的提取器使用冻结的1280维投影AV-HuBERT特征、目标条件训练和分层特征调制,通过Qwen3-ASR-1.7B字符错误率(CER)评估内容,通过WeSpeaker ResNet34加重叠语音检测(OSD)评估目标身份。在完整清单上,最佳存档检查点获得0.2261 CER、82.22%严格输出正确率、69.53%双输出严格成功率。

英文摘要

Audio-visual target speaker extraction should return the speaker indicated by the video, yet a separator can ignore the visual cue and repeatedly output the acoustically dominant voice. We introduce REAL-2MIX, a Mandarin AV-TSE benchmark of jointly recorded real two-speaker mixtures with synchronized multi-view video. Each scene also contains preceding A-only and B-only stages that provide in-scene speaker references. It contains 77 scenes and 7,598 clips (11.84 hours), including 6,042 dual-annotated mixtures. After processing there leave 6,038 evaluable mixtures and 12,076 target-speaker rows. We additionally curate VOXBLINK2-AVSE from VoxBlink2, comprising 250,828 synchronized audio--lip-ROI pairs from 28,421 identities and 766.17 hours of speech. Our extractor uses frozen, 1,280-dimensional projected AV-HuBERT features, target-conditioned training, and layer-wise feature modulation. We jointly evaluate content with Qwen3-ASR-1.7B CER and target identity with WeSpeaker ResNet34 plus Overlapped Speech Detection (OSD). On the complete manifest, the best archived checkpoint obtains 0.2261 CER, 82.22% strict output correctness, and 69.53% both-output strict success.

↑