arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向可操作的对话智能:一种语音智能框架

Towards Operational Conversational Intelligence: A Speech Intelligence Framework

C. Vishnoi, S. Khurana, A. Timmapur, S. Rai, S. Mohanty

arXiv 2607.24958首次发表:更新:

发表机构

Indian Institute of Technology Kanpur; EXL(印度理工学院坎普尔分校; EXL公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对佩戴式摄像机音频处理难题,提出双路径对话智能框架,含说话者分离与转录分支,经融合输出实现单词级说话者归因。在特定数据集上评估,结果显示其能改善相关任务,为未来系统提供可扩展基础。

AI 中文摘要

佩戴式摄像机(BWC)音频存在独特挑战,如高环境噪声、可变录制条件和多个重叠说话者,这使自动转录和说话者标注具有挑战性。我们提出了一种双路径对话智能框架,对原始BWC音频进行预处理,将处理管道分为说话者分离分支和自动语音识别(ASR)分支,并融合其输出。说话者分离分支使用去噪前端、语音活动检测和带有TitaNet嵌入的NVIDIA多尺度说话者分离解码器。转录分支使用响度归一化和带有强制对齐及概率引导语音分割的WhisperX(Large-v3)。最后通过将每个识别出的单词分配到时间重叠最大的说话者片段来进行单词级说话者归因。我们在一个由美国和英国公开的警察佩戴式摄像机录音构建的精选数据集上评估了该框架。实验结果表明,特定任务的声学调节和概率引导语音分割在具有挑战性的佩戴式摄像机录制条件下改善了说话者分离、转录和单词级说话者归因。所提出的模块化架构为未来的说话者感知对话智能系统提供了可扩展基础。

英文摘要

Body-worn camera (BWC) audio presents unique challenges including high ambient noise, variable recording conditions, and multiple overlapping speakers that make automated transcription and speaker labeling challenging. We propose a dual-path conversational intelligence framework that preprocesses raw BWC audio, separates the processing pipeline into a diarization branch and an ASR branch, and fuses their outputs. The diarization branch uses a denoising front-end (DeepFilterNet), voice activity detection (VAD), and NVIDIA's Multi-Scale Speaker Diarization Decoder (MSDD) with TitaNet embeddings. The transcription branch uses loudness normalization and WhisperX (Large-v3) with forced alignment and probability-guided speech segmentation. Finally, word-level speaker attribution is performed by assigning each recognized word to the speaker segment with the greatest temporal overlap. We evaluate the proposed framework on a curated body-worn camera dataset constructed from publicly available U.S. and U.K. police body-worn camera recordings. Experimental results demonstrate that task-specific acoustic conditioning and probability-guided speech segmentation improve speaker diarization, transcription, and word-level speaker attribution under challenging body-worn camera recording conditions. The proposed modular architecture provides an extensible foundation for future speaker-aware conversational intelligence systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑