用于2026年智能眼镜挑战赛TSA-ASR任务的tttAI系统
The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026
浏览论文内容
中文总结 AI 辅助
介绍用于2026年智能眼镜挑战赛TSA-ASR任务的tttAI系统,针对长音频、多说话人及语音重叠挑战,提出级联架构,含说话人日记化等模块,最终系统用特定模型和通道,在赛道2获34.04%的tcpCER并排名第二。
中文摘要 AI 辅助
本文介绍了提交至2026年智能眼镜挑战赛TSA-ASR任务的tttAI系统,该系统在两人对话(赛道1)和多方会议(赛道2)上进行评估。任务要求从智能眼镜录音中进行带时间戳的说话人归因语音识别,因音频长、说话人多和语音频繁重叠极具挑战。我们提出了一种级联架构,包括说话人日记化、重叠检测、目标说话人提取、后处理和自动语音识别。日记化模块通过WavLM-Large提取特征,用Conformer编码器进行逐帧说话人分类,再通过嵌入聚类生成全局说话人片段。对于重叠区域,应用基于WeSep的目标说话人提取模型及ECAPA-TDNN说话人嵌入。提取不可靠时采用主导说话人回退策略。最终系统使用FireRedASR2-AED及第一麦克风通道。提交系统总参数约15.3亿。在赛道1上,系统tcpCER为7.10%。在赛道2上,tcpCER为34.04%,在排行榜上排名第二。
英文摘要
This paper presents the tttAI system submitted to the TSA-ASR task of the SmartGlasses Challenge 2026, evaluated on both two-person dialogues (Track 1) and multi-party meetings (Track 2). The task requires time-stamped speaker-attributed speech recognition from smart-glasses recordings. This is particularly challenging due to long-form audio, multiple speakers, and frequent overlapping speech. We proposed a cascaded architecture consisting of speaker diarization, overlap detection, target-speaker extraction, post-processing, and automatic speech recognition. The diarization module extracts features via WavLM-Large, performs frame-wise speaker classification with a Conformer encoder, and then generates global speaker segments through embedding clustering. For overlapped regions, we apply a WeSep-based target-speaker extraction model with ECAPA-TDNN speaker embeddings. When the extraction is unreliable, a dominant-speaker fallback strategy is used. The final system uses FireRedASR2-AED with the first microphone channel. The submitted system has a total parameter count of approximately 1.53B. On Track 1, our system achieves a tcpCER of 7.10%. On Track 2, it achieves a tcpCER of 34.04% and ranks second on the leaderboard.