arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17867eess.AS

用于2026年智能眼镜挑战赛TSA-ASR任务的tttAI系统

The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026

Xuanji He, Gaoyang Dong, Xiaoxiao Li, Minchuan Chen, Fengjie Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

介绍用于2026年智能眼镜挑战赛TSA-ASR任务的tttAI系统,针对长音频、多说话人及语音重叠挑战,提出级联架构,含说话人日记化等模块,最终系统用特定模型和通道,在赛道2获34.04%的tcpCER并排名第二。

中文摘要 AI 辅助

本文介绍了提交至2026年智能眼镜挑战赛TSA-ASR任务的tttAI系统,该系统在两人对话(赛道1)和多方会议(赛道2)上进行评估。任务要求从智能眼镜录音中进行带时间戳的说话人归因语音识别,因音频长、说话人多和语音频繁重叠极具挑战。我们提出了一种级联架构,包括说话人日记化、重叠检测、目标说话人提取、后处理和自动语音识别。日记化模块通过WavLM-Large提取特征,用Conformer编码器进行逐帧说话人分类,再通过嵌入聚类生成全局说话人片段。对于重叠区域,应用基于WeSep的目标说话人提取模型及ECAPA-TDNN说话人嵌入。提取不可靠时采用主导说话人回退策略。最终系统使用FireRedASR2-AED及第一麦克风通道。提交系统总参数约15.3亿。在赛道1上,系统tcpCER为7.10%。在赛道2上,tcpCER为34.04%,在排行榜上排名第二。

英文摘要

This paper presents the tttAI system submitted to the TSA-ASR task of the SmartGlasses Challenge 2026, evaluated on both two-person dialogues (Track 1) and multi-party meetings (Track 2). The task requires time-stamped speaker-attributed speech recognition from smart-glasses recordings. This is particularly challenging due to long-form audio, multiple speakers, and frequent overlapping speech. We proposed a cascaded architecture consisting of speaker diarization, overlap detection, target-speaker extraction, post-processing, and automatic speech recognition. The diarization module extracts features via WavLM-Large, performs frame-wise speaker classification with a Conformer encoder, and then generates global speaker segments through embedding clustering. For overlapped regions, we apply a WeSep-based target-speaker extraction model with ECAPA-TDNN speaker embeddings. When the extraction is unreliable, a dominant-speaker fallback strategy is used. The final system uses FireRedASR2-AED with the first microphone channel. The submitted system has a total parameter count of approximately 1.53B. On Track 1, our system achieves a tcpCER of 7.10%. On Track 2, it achieves a tcpCER of 34.04% and ranks second on the leaderboard.

↑