arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.03670eess.ASeess.SP

CHILDES对齐:通过多模型时间戳集成的精心策划儿童语音数据集

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

发表机构伊利诺伊大学厄巴纳-香槟分校 · 香港中文大学(深圳) · IBM研究院
另 1 家 · 查看机构详情
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • IBM Research(IBM研究院)
  • University at Buffalo(布法罗大学)

机构由 AI 辅助整理,请以论文原文为准。

Haolong Zheng, Yuanzhuo Hu, Xinyu Liang, Vishal Sunder, Dancheng Liu, Jinjun Xiong, Samuel Thomas, Brian Kingsbury, Zhizheng Wu, Mark A. Hasegawa-Johnson

首次发表
浏览论文内容

中文总结 AI 辅助

研究儿童语音数据集时间戳问题,提出BEACON框架,通过聚合多ASR模型知识优化时间戳,贡献经整理含正确时间戳的数据集及子集,微调后在基准测试中WER降低。

中文摘要 AI 辅助

CHILDES是大规模儿童语音语料库,但话语级时间戳有问题。本文提出BEACON框架,通过聚合多现成ASR模型知识优化时间戳,该框架与语料库无关,适用于时间戳不可靠或缺失的情况。利用此框架整理发布含正确时间戳的数据集及子集,微调后在基准测试中WER降低。

英文摘要

CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and language development. However, utterance-level timestamps provided in this corpus are often noisy, incomplete, or misaligned with the audio. As a result, utterances cannot always be reliably localized within long recordings, which limits the direct use of these data for training and evaluating speech models. In this work, we propose BEACON (Boundary Estimation via Alignment CONsensus), an ensemble timestamp-curation framework that refines utterance-level timestamps by aggregating knowledge from multiple off-the-shelf ASR models. Specifically, each model's word-level timestamp predictions are first aligned to provided human transcripts, and the final utterance time boundaries are determined by a consensus voting strategy. The framework is corpus-agnostic and applies to any long-form recording paired with a trusted transcript whose timestamps are unreliable or missing, offering a general recipe for timestamp curation. Leveraging this pipeline, we curate and release a 413-hour general-purpose child-speech dataset with corrected utterance-level timestamps, together with a 283-hour quality-controlled subset for ASR training. Fine-tuning on this subset yields up to an average 19.5% relative WER reduction on four out-of-domain child-speech benchmarks.

↑