AI 中文总结
研究提出用大语言模型和文本转语音技术生成基于对话课程的半自动系统,经准实验探索其教育潜力。系统通过三阶段工作流程增强教育工作者,引入新方法。实验表明对话TTS在理解等方面优于单声道TTS,为TTS音频教育可接受性及课程形式设计提供依据。
AI 中文摘要
本研究提出了一种使用大语言模型(LLMs)和文本转语音(TTS)技术生成基于对话的课程的半自动系统,并通过实际的准实验探索性地研究其教育潜力。该系统通过三阶段的人工参与工作流程(基于LLM的幻灯片/旁白生成、教育工作者审查、自动视听整合)来增强而非取代教育工作者,并引入了一种基于认知学徒理论生成专家-新手对话旁白的新方法。在一项对245名高一学生的研究中,他们依次体验了三种课程形式(教师语音、单声道TTS、对话TTS;各阶段内容不同,限制了形式/内容分离),我们进行了主体内(Friedman检验,N<=183)和重复横截面(Mann-Whitney U,N=229/206)分析。TOST等效性测试表明,与教师语音相比,TTS音频并未显著降低学习体验。对话TTS在理解(p=.006,q=.025)和认知参与(p=.019,q=.048)方面显著优于单声道TTS;经FDR校正后,愉悦感无显著差异(q=.081),但在控制先验知识后达到显著水平(比例优势模型,OR=1.65,q=.025),且这些优势并非归因于先验知识不平衡。相反,单声道TTS在音频自然度方面更优(p<.001,q<.001,r=-.238),揭示了对话的益处与更高的额外认知负荷之间的权衡。66.9%的学习者更喜欢对话形式,认为其最有趣(p<.001)。这些结果反映了固定顺序设计;在将其推广为课程形式的效果之前,需要进行重复实验。本研究为TTS音频的教育可接受性和TTS课程形式设计提供了理论和实证基础。
英文摘要
This study proposes a semi-automated system for generating dialogue-based lessons using Large Language Models (LLMs) and Text-to-Speech (TTS) technology, and exploratorily examines its educational potential via a practical quasi-experiment. The system augments rather than replaces educators through a three-stage human-in-the-loop workflow (LLM-based slide/narration generation, educator review, automated audiovisual integration), and introduces a novel method for generating Expert-Novice dialogue narration based on cognitive apprenticeship theory. In a study of 245 first-year high school students who sequentially experienced three lesson formats (instructor voice, single-speaker TTS, dialogue TTS; content differed across sessions, limiting format/content separation), we conducted within-subject (Friedman test, N<=183) and repeated cross-sectional (Mann-Whitney U, N=229/206) analyses. TTS audio did not substantially degrade the learning experience versus instructor voice, supported by TOST equivalence testing. Dialogue TTS was significantly superior to single TTS in comprehension (p=.006, q=.025) and cognitive engagement (p=.019, q=.048); enjoyment was non-significant after FDR correction (q=.081) but reached significance after controlling for prior knowledge (proportional-odds model, OR=1.65, q=.025), and these advantages were not attributable to prior-knowledge imbalance. Conversely, single TTS was superior in audio naturalness (p<.001, q<.001, r=-.238), revealing a trade-off between dialogue's benefits and higher extraneous cognitive load. Dialogue format was preferred by 66.9% of learners as most enjoyable (p<.001). These results reflect a fixed-order design; replication is needed before generalizing them as effects of lesson format. This study provides a theoretical and empirical basis for the educational acceptability of TTS audio and for TTS lesson-format design.