arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

美国网络流录音的宗教广播文字转写大规模语料库

A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

Samuel Bestvater, Athena Chapekis, Skyler Seets, Anna Lieb, Sono Shah, Aaron Smith

arXiv 2607.26249首次发表:更新:

发表机构

Pew Research Center(皮尤研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建了2025年7月美国宗教广播网络流录音的大规模文字转写语料库,可用于宗教传播、社会议题分析及语音处理等相关领域研究。

AI 中文摘要

宗教广播是美国一种广泛存在但研究不足的大众传播形式,其内容层面的分析因缺乏大规模文字转写数据而受限。本数据描述符呈现了一个语料库,该语料库包含2025年7月一个月内从实时网络流中捕获的英语宗教广播文字转写内容。研究按滚动计划对785个不同流的15分钟片段进行录音,这些流共同转播了2000多个调幅(AM)和调频(FM)电台的信号,共产生超过70万条录音和6000多万条带说话人标注的语音行。每条录音都通过自动化流程进行文字转写和说话人标注,并使用大语言模型按节目类型和主题进行分段与标注。该语料库以流元数据、录音元数据和文字转写行的关联表形式组织,支持对跨地区和传统的宗教广播的描述性研究、对宗教媒体中社会与政治议题讨论方式的分析,以及对代表性不足领域的语音处理研究。

英文摘要

Religious radio is a widespread but understudied form of mass communication in the United States, and content-level analysis of it has been constrained by the absence of large-scale transcript data. This Data Descriptor presents a corpus of transcribed English-language religious radio broadcasts captured from live webstreams over a one-month period in July 2025. Fifteen-minute segments were recorded on a rolling schedule from 785 distinct streams, which together rebroadcast the signals of more than two thousand AM and FM stations, yielding over 700,000 recordings and more than 60 million diarized lines of speech. Each recording was transcribed and speaker-diarized with an automated pipeline, and segmented and labeled by programming format and topic using a large language model. The corpus is organized as linked tables of stream metadata, recording metadata, and transcript lines. It supports descriptive study of religious broadcasting across regions and traditions, analysis of how social and political issues are discussed in religious media, and speech-processing research in an underrepresented domain.

CommentsPresented at IC2S2 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑