arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

第二语言朗读语音中的声学间隙位置

Acoustic gap placement in second-language read speech production

Peyman Jahanbin

arXiv 2610.02582首次发表:更新:

发表机构

University of Kansas(堪萨斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过分析同一朗读段落中词间声学间隙的位置,发现普通话母语者比美国英语母语者在非标点位置产生更多停顿,表明间隙位置可作为衡量第二语言口语流利性的有效维度。

AI 中文摘要

说话者之间的差异不仅在于他们产生的静音时长,还在于中断发生的位置,而这些信息被全局停顿计数所掩盖。我们分析了语音口音档案库中115个公开可用的同一朗读段落录音:其中57名说话者的档案元数据将其第一语言列为普通话,且出生地为中国大陆;另有58名美国英语第一语言说话者,出生于美国中西部。使用单词级强制对齐来提取词间声学间隙,并将每个位置分类为标点标记或非标点位置。在250毫秒阈值下,标点标记的间隙在普通话组和英语组之间相似(平均值分别为5.56和5.47),而非标点位置的间隙在普通话组中更为频繁(平均值分别为3.67和0.43;调整后的比率比为11.98,95%置信区间为6.53-22.00)。一个具有说话者和段落位置随机截距的交叉效应逻辑模型证实,这种对比不成比例地集中在非标点位置(交互作用比值比为9.44,95%置信区间为5.23-17.02)。在500毫秒阈值下,在移除标点相邻位置后,在排除严重转录偏差案例后,以及当分析仅限于声学确认的静音时,该模式仍然存在。普通话组近一半的非标点间隙发生在探索性单编码器注释中归类为短语内部的位置。在这个固定段落语料库中,相对于文本和句法结构的声学间隙位置比标点标记的停顿更能清晰地区分两组。结果支持将位置敏感的测量作为流畅性障碍的一个可复现维度,但并未建立因果机制、熟练度差异或感知后果。

英文摘要

Speakers can differ not only in how much silence they produce but also in where interruptions fall, information that global pause counts obscure. We analyzed 115 publicly available Speech Accent Archive recordings of the same read passage: 57 speakers whose archive metadata listed Mandarin as their first language and a mainland-China birthplace, and 58 American English first-language speakers born in the U.S. Midwest. Word-level forced alignment was used to extract interword acoustic gaps and classify each position as punctuation-marked or unpunctuated. At a 250 ms threshold, punctuation-marked gaps were similar for the Mandarin and English groups (means 5.56 and 5.47), whereas unpunctuated-position gaps were more frequent in the Mandarin group (means 3.67 and 0.43; adjusted rate ratio 11.98, 95% confidence interval 6.53-22.00). A crossed-effects logistic model with random intercepts for speakers and passage positions confirmed that the contrast was disproportionately concentrated at unpunctuated positions (interaction odds ratio 9.44, 95% confidence interval 5.23-17.02). The pattern persisted at a 500 ms threshold, after removing positions adjacent to punctuation, after excluding severe transcript-deviation cases, and when analysis was restricted to acoustically confirmed silence. Nearly half of the Mandarin group's unpunctuated gaps occurred at positions classified as within-phrase in an exploratory single-coder annotation. In this fixed-passage corpus, acoustic gap placement relative to textual and syntactic structure distinguished the groups more clearly than punctuation-marked pausing. The results support placement-sensitive measurement as a reproducible dimension of breakdown fluency while not establishing a causal mechanism, proficiency difference, or perceptual consequence.

Comments24 pages, appendix included, under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑