arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

公共话语语料库(PDC):面向效价与认知情态且含目标说话人参与的说话人标注数据集

The Public Discourse Corpus (PDC): A Speaker-Attributed Dataset for Valence and Epistemic Modality with Target Speaker Participation

Bo Chen

arXiv 2609.20232首次发表:更新:

发表机构

Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出公共话语语料库(PDC),首个联合标注效价与认知情态的公众人物访谈数据集,含998个视频、18.6万句,并引入目标说话人参与(TSP)分类体系及音频优先二值化流程,提升语料质量与可复用性。

AI 中文摘要

我们推出了公共话语语料库(PDC),这是首个对公众人物访谈语音进行情感效价和认知情态联合标注的数据集。该语料库包含来自七个专业领域100位说话人的998个视频,经过句子分割和过滤后,共获得186,642个句子(310万词)。为确保所有保留的视频包含目标说话人可分析的语音,我们引入了目标说话人参与(TSP)——一个五类标注分类体系,并记录了标注者间信度(κ = 0.616)——作为任何语料库构建项目均可采用的关键方法论贡献。通过结合本地Whisper ASR与pyannote说话人分离的音频优先二值化流程,将目标说话人的话轮与采访者及第三方语音区分开来,该流程已作为开源实现发布。我们发布了标注语料库、标注工具、跨提供者验证样本以及完整处理流程。数据集可在该https URL获取。

英文摘要

We introduce the \textbf{Public Discourse Corpus (PDC)}, the first dataset of public-figure interview speech jointly annotated for affective valence and epistemic modality. The corpus contains 998 videos from 100 speakers across seven professional domains, yielding 186,642 sentences (3.1 million words) after sentence segmentation and filtering. To ensure that all retained videos contain analyzable speech from the intended speaker, we introduce \textbf{Target Speaker Participation (TSP)}---a five-category annotation taxonomy with documented inter-annotator reliability ($κ= 0.616$)---as a key methodological contribution that any corpus construction project can adopt. Target-speaker turns are separated from interviewer and third-party speech through an \textbf{audio-first diarization pipeline} combining local Whisper ASR with pyannote speaker separation, released as an open-source implementation. We release the annotated corpus, the annotation tools, the cross-provider validation sample, and the complete processing pipeline. The dataset is available at https://huggingface.co/datasets/ictchenbo/public-discourse-corpus.

Comments16 pages, 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑