语音语言模型用于全场会议说话人日志:能力与局限
Speech Language Models for Full-Meeting Speaker Diarization: Capabilities and Limitations
浏览论文内容
中文总结 AI 辅助
本研究使用ESPnet-SpeechLM作为骨干,将全场会议说话人日志建模为结构化令牌的自回归生成,比较事件与帧表示,发现事件表示更稳定,且需结合显式说话人跟踪与重叠感知生成以提升性能。
中文摘要 AI 辅助
语音语言模型(SpeechLMs)的最新进展,将大语言模型与语音基础模型相结合,实现了语音处理任务的统一序列建模。然而,许多基于SpeechLM的说话人日志(SD)方法紧密耦合自动语音识别(ASR),并使用词级指标进行评估,这使得难以独立于ASR准确性来评估SD性能。在本工作中,我们研究ESPnet-SpeechLM作为基于令牌的骨干网络,用于生成SD假设,将SD表述为以声学输入为条件的结构化令牌的自回归生成。我们系统比较了两种输出表示:一种是基于事件的表示,显式建模说话人轮次的开始和结束时间戳;另一种是基于帧的表示,预测帧级说话人活动。为了提供结构化的对话线索,我们进一步在输出序列中纳入辅助任务,包括语音活动检测、重叠语音检测和说话人轮次计数。在多个会议数据集上,我们发现基于事件的表示比基于帧的表示产生更稳定和一致的SD输出。我们的分析表明,SpeechLM生成的输出编码了有用的时间SD结构,但全场会议SD仍受限于录音级说话人跟踪和重叠相关的遗漏。显式的说话人链接后处理显著减少了说话人混淆,这表明基于SpeechLM的稳健SD需要持久的说话人跟踪和重叠感知生成。
英文摘要
Recent advances in Speech Language Models (SpeechLMs), which integrate large language models with speech foundation models, have enabled unified sequence modeling of speech processing tasks. However, many SpeechLM-based approaches to speaker diarization (SD) are tightly coupled with automatic speech recognition (ASR) and evaluated using word-level metrics, making it difficult to assess SD performance independent of ASR accuracy. In this work, we investigate ESPnet-SpeechLM as a token-based backbone for generating SD hypotheses, formulating SD as autoregressive generation of structured tokens conditioned on acoustic input. We systematically compare two output representations: an event-based representation that explicitly models speaker turn onset and offset timestamps, and a frame-based representation that predicts frame-level speaker activity. To provide structured conversational cues, we further incorporate auxiliary tasks including speech activity detection, overlapped speech detection, and speaker turn counting within the output sequence. Across multiple meeting datasets, we find that event-based representations produce more stable and consistent SD outputs than frame-based representations. Our analysis shows that outputs generated by SpeechLMs encode useful temporal SD structure, but full-meeting SD remains limited by recording-level speaker tracking and overlap-related misses. Explicit speaker-linking post-processing substantially reduces speaker confusion, suggesting that robust SpeechLM-based SD requires persistent speaker tracking and overlap-aware generation.
发表机构
- College of Information Science, University of Arizona(亚利桑那大学信息科学学院)
- Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所)
机构由 AI 辅助整理,请以论文原文为准。