AI 中文总结
本研究提出含三个系统的AI引导式学习框架,通过深度学习音视频处理技术,解决学习中长内容耗时、技能模仿反馈不足的问题,相关系统在播放效率、视频时长缩减、发音提升上均有良好效果。
AI 中文摘要
音视频已成为主要的学习媒介,但学习者面临两个长期存在的挑战:连续消费长内容的时间成本,以及基于模仿的技能获取缺乏可扩展的反馈。本论文提出一种AI引导式学习框架,支持三个相互关联的阶段:消费、理解和模仿。该框架开发并评估了三个系统:AIxSpeed利用语音识别模型的置信度作为听力难度的代理,在音素级别动态调整音频播放速度;FastPerson生成保留视觉和听觉信息的多模态视频摘要,允许学习者按章节在摘要版和完整版之间切换;Profy从大量未标注的语音数据中学习熟练度,并可视化分类器相关区域和模型衍生的声学距离,以支持发音练习。技术和用户评估显示:AIxSpeed在LibriSpeech上实现平均播放倍率1.30x,在UME-ERJ上实现1.29x,且获得的平均意见得分高于匹配的恒速播放;FastPerson将观看时间减少53%,与正常播放相比测验得分无统计学显著差异;Profy显示发音可理解性有观察到的提升,练习前后的置信区间无重叠。这些系统共同证明了深度学习如何在支持高效内容消费、多模态理解和重复技能练习的同时,让学习者能够访问原始材料。
英文摘要
Audio and video have become major learning media, but learners face two persistent challenges: the time cost of consuming long-form content sequentially and the lack of scalable feedback for imitation-based skill acquisition. This dissertation proposes an AI-guided learning framework that supports three interconnected stages: Consume, Understand, and Imitate. It develops and evaluates three systems. AIxSpeed dynamically adjusts audio playback speed at the phoneme level using speech-recognition-model confidence as a proxy for listening difficulty. FastPerson generates multimodal video summaries that preserve visual and auditory information and lets learners switch between summarized and full versions by chapter. Profy learns proficiency from largely unannotated speech data and visualizes classifier-relevant regions and model-derived acoustic distances to support pronunciation practice. Technical and user evaluations show that AIxSpeed achieved average playback factors of 1.30x on LibriSpeech and 1.29x on UME-ERJ and received higher mean opinion scores than matched constant-speed playback; FastPerson reduced viewing time by 53% with no statistically significant difference in quiz scores compared with normal playback; and Profy showed an observed improvement in pronunciation intelligibility, with non-overlapping pre- and post-practice confidence intervals. Together, these systems demonstrate how deep learning can support efficient content consumption, multimodal understanding, and repeated skill practice while retaining learner access to the original material.
Comments118 pages, 26 figures, 4 tables. Doctoral dissertation, Doctor of Interdisciplinary Informatics, The University of Tokyo; degree awarded September 19, 2025