arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自发对话中标准化语音单元标注的自动化流程

An automated pipeline for standardised speech-unit annotation in spontaneous dialogue

Hanlu He, Harald Vilhelm Skat-Rørdam, Ingvi Örnólfsson, Ivana Konvalinka

arXiv 2610.03078首次发表:更新:

发表机构

Technical University of Denmark(丹麦技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该论文提出一种自动化流程,用于从自发对话录音中提取轮次和反馈语,结合多种技术,在丹麦语语料上达到F1=0.621,支持半自动标注工作流。

AI 中文摘要

量化对话动态需要可靠地识别交互单元及其时间边界,但仅凭语音活动无法区分对话轮次与听者反馈或轮内停顿。我们提出了一种自动化流程,用于从自发二元对话的分声道录音中提取轮次和反馈语,旨在为后续人工审核提供一致的首轮标注。该流程结合了语音活动检测、声道能量过滤、时间合并、自动语音识别和基于上下文的后处理。我们使用片段级检测可靠性和时间边界误差,在来自33个二人组的99段十分钟丹麦语对话上评估了该流程。对话在正常和不对称听力条件下录制;后者通过骨传导耳机向一名参与者传递语音形噪声。总体检测可靠性为F1=0.621,轮次(F1=0.624)和反馈语(F1=0.618)的性能相似。对于成功匹配的片段,轮次的中位绝对起始和结束误差分别为0.150秒和0.160秒,反馈语分别为0.130秒和0.180秒。轮次边界的平均误差显著更大,表明存在少量大的边界错配。两种实验听力条件下的性能无显著差异。在四段对话的案例研究中,流程与人工的一致性低于且变异性高于人工标注者之间的一致性,并随参数设置而变化。这些结果支持该流程作为半自动标注工作流中的自动化首轮,为更标准化和可复现的对话动态标注提供一致基础。

英文摘要

Quantifying conversational dynamics requires reliable identification of interactional units and their temporal boundaries, but speech activity alone does not distinguish conversational turns from listener feedback or within-turn pauses. We present an automated pipeline for extracting turns and backchannels from separate-channel recordings of spontaneous dyadic conversation, designed to provide a consistent first-pass annotation for subsequent human review. The pipeline combines voice activity detection, channel-energy filtering, temporal merging, automatic speech recognition, and context-based post-processing. We evaluated the pipeline on 99 ten-minute Danish conversations from 33 dyads using segment-level detection reliability and temporal boundary error. Conversations were recorded under both normal and asymmetric listening conditions. In the latter, speech-shaped noise was delivered to one participant through bone-conduction headphones. Overall detection reliability was F1=0.621, with similar performance for turns F1=0.624 and backchannels F1=0.618. For successfully matched segments, median absolute onset and offset errors were 0.150 and 0.160s for turns and 0.130 and 0.180s for backchannels, respectively. Mean errors were substantially larger for turn boundaries, indicating a smaller number of large boundary mismatches. Performance did not differ significantly across the two experimental listening conditions. In a four-conversation case study, pipeline-human agreement was lower and more variable than human inter-annotator agreement and varied across parameter settings. These results support the pipeline as an automated first pass within a semi-automated annotation workflow, providing a consistent basis for more standardised and reproducible annotation of conversational dynamics.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑