arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03952cs.AI

TACT:面向教学自适应英语辅导的分类对齐后训练

TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring

Dongjie Yang, Siyan Lin, Leixian Shen, Rui Sheng, Huamin Qu, Zixin Chen

AI总结:

本研究提出基于人类研究的TACT框架,构建含双分类体系的数据集后,通过特定训练得到TACTutor,在自研基准及盲法研究中表现优于基线,为自适应ESL辅导器开发提供开放基础。

AI中文摘要:

大型语言模型(LLM)正越来越多地用于为英语作为第二语言(ESL)学习者提供对话练习。然而,有效的ESL辅导不仅需要生成流畅的回复,还要求辅导者根据学习者的行为和对话语境选择合适的教学行动。人类辅导研究提供了自适应支持的原则,但这些原则通常是针对特定任务的,且未能充分整合到基于LLM的ESL辅导器的训练与评估中。我们提出TACT(Taxonomy-Aligned Conversational Tutor,分类对齐对话辅导器),这是一个基于人类研究的框架,用于教学自适应ESL辅导器的后训练与评估。我们参考已有文献,开发了两个互补的分类体系:包含13种辅导者回复策略的辅导者策略分类体系,以及通过行动类型和状态表征学习者行为的学习者行动分类体系。基于这些分类体系,我们构建了TACTCorpus,该数据集为260段真实的师生对话补充了32379条标注和质量控制的增强训练数据。随后,我们通过监督微调再加上分类对齐的分组相对策略优化,对Qwen3.5-4B进行后训练,得到TACTutor,使其优化目标不仅限于参考模仿,更聚焦于支架式教学质量。在由78个真实辅导语境构成的、策略均衡的诊断基准TACTBench上,TACTutor相比其骨干模型性能提升了20.30%,并在相同协议下优于所有评估的专有基线模型,同时在已有的外部教育基准上保持了骨干模型的性能;在一项有50名学习者参与的盲法研究中,它还获得了被评估辅导器中最高的总体平均评分。我们发布了该数据、基准和模型权重,为开发教学自适应ESL辅导器提供了开放基础。

英文摘要:

Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-language (ESL) learners. Effective ESL tutoring, however, requires more than fluent response generation: a tutor must select an appropriate pedagogical action based on learner behavior and dialogue context. Human-tutoring research offers principles for adaptive support, but they are often task-specific and remain insufficiently integrated into LLM-based ESL tutor training and evaluation. We present TACT (Taxonomy-Aligned Conversational Tutor), a human-grounded framework for post-training and evaluating pedagogically adaptive ESL tutors. Drawing on established literature, we develop two complementary taxonomies: the Tutor-Strategy Taxonomy with 13 tutor response strategies and the Student-Move Taxonomy characterizing learner behavior by move type and status. Using these taxonomies, we construct TACTCorpus, which enriches 260 authentic teacher-student conversations with 32,379 annotations and quality-controlled augmented training data. We then post-train Qwen3.5-4B through supervised fine-tuning followed by taxonomy-aligned Group Relative Policy Optimization, producing TACTutor and optimizing it for scaffolding quality rather than reference imitation alone. On TACTBench, a strategy-balanced diagnostic benchmark comprising 78 authentic tutoring contexts, TACTutor improves over its backbone by 20.30% and outperforms all evaluated proprietary baselines under the same protocol, while maintaining backbone performance on established external educational benchmarks; in a blinded study with 50 learners, it also receives the highest overall mean rating among the evaluated tutors. We release the data, benchmark, and model weights, providing an open foundation for developing pedagogically adaptive ESL tutors.

↑