AI 中文总结
该研究评估了LLM替代语言教师提供教学反馈与解释的能力,发现其纠错能力出色但教学解释不足,AI辅助语言学习热情或超出对其教学能力的认知。
AI 中文摘要
尽管各类机构如今积极鼓励课堂中使用大型语言模型(LLM),但我们仍缺乏对这些模型完成语言教学核心任务表现的严谨、系统评估。本文探究最先进的大型语言模型能否提供语言学习者所需的纠错反馈与教学法解释。该研究测试了多个大型语言模型识别、纠正并解释英语常见学习者错误的能力,通过系统调整模型参数,探究这些技术调整如何影响输出质量、教学清晰度与一致性,同时使用检索增强生成(RAG)查询教学法数据。评估采用自动指标(GLEU、BERTScore),也采用人类专家判断,以捕捉纯计算指标遗漏的维度:语言细微差别、文化敏感性与教学适宜性。尽管模型展现出出色的表层纠错能力,但其解释往往缺乏有效语言教学所需的术语及领域知识,表明当前对AI辅助语言学习的热情可能已超出我们对这些系统实际教学能力的理解。
英文摘要
While various organizations now actively encourage LLM use in classrooms, we still lack rigorous, systematic evaluations of how well these models actually perform the fundamental tasks of language pedagogy. This paper examines whether state-of-the-art LLMs can deliver the kind of corrective feedback and methodological explanations that language learners need. The study tests multiple large language models on their ability to identify, correct, and explain common learner mistakes in English, by systematically varying model parameters to investigate how these technical adjustments affect output quality, pedagogical clarity, and consistency, along with using retrieval-augmented generation to query methodological data. The evaluation employs automated metrics (GLEU, BERTScore) but also human expert judgments to capture dimensions that purely computational measures miss: linguistic nuance, cultural sensitivity, and instructional appropriateness. While models demonstrate impressive surface-level correction abilities, their explanations often lack the terminological and domain knowledge that effective language teaching requires, suggesting that current enthusiasm for AI-assisted language learning may be outpacing our understanding of these systems' actual pedagogical competence.
Journal refOxford Intersections: AI in Society (Oxford, online edn, Oxford Academic, 20 Mar. 2025 - )
DOI:10.1093/9780198945215.003.0252