arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29804cs.MA

REAT:一种用于多轮数学教学的反思性经验增强辅导框架

REAT: A Reflective Experience-Augmented Tutoring Framework for Multi-turn Mathematical Instruction

  • Nanyang Technological University(南洋理工大学)
  • Zhejiang Normal University(浙江师范大学)
  • Telfer School of Management, University of Ottawa(渥太华大学特尔弗管理学院)
  • University of Calabria(卡拉布里亚大学)
  • APSS, Hong Kong Polytechnic University(香港理工大学应用物理及材料科学系)
  • School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
  • Squirrel Ai Learning(松鼠AI)

机构由 AI 辅助整理,请以论文原文为准。

Jianheng Zhou, Chaoli Zhang, Xingjun Wei, Xinliang Zhou, Giancarlo Fortino, Xing Fan, Yanfeng Wang, Qingsong Wen, Haoyang Li

AI总结:

提出REAT框架,通过多智能体提炼历史对话经验并实时检索注入,显著提升多轮数学辅导效果,尤其在复杂低分场景,且经验泛化性强。

AI中文摘要:

当前的大型语言模型(LLMs)擅长解决复杂的数学问题,但这种熟练度并不天然转化为有效的教学辅导。虽然先进的LLM辅导者可能利用多智能体框架或微调技术,但大多数仍缺乏一种机制来系统地积累和复用教学经验,从而限制了它们在流畅的多轮交互中适应多样化学生需求的能力。为弥补这一差距,我们提出了反思性经验增强辅导(REAT)框架,该框架将历史对话中的经验提炼与实时自适应检索相结合。由多智能体观察者-批评者-导师(OCM)提炼流程驱动,REAT回顾过去的对话轨迹,并将原始交互提炼为结构化的、与问题无关的教学经验。在实时辅导中,状态感知检索模块注入这些精选经验,以根据学生的认知状态提供自适应脚手架。实验表明,所提出的框架显著优于仅提示和监督微调(SFT)基线,特别是在改善复杂、低分辅导场景方面。至关重要的是,提炼出的经验在不同模型架构和数学数据集上展现出强大的泛化能力。

英文摘要:

Current Large Language Models (LLMs) excel at solving complex mathematical problems, yet this proficiency does not inherently translate into effective tutoring. While advanced LLM tutors may leverage multi-agent frameworks or fine-tuning, most still lack a mechanism to systematically accumulate and reuse pedagogical experience over time, limiting their adaptability to diverse student needs during fluid, multi-turn interactions. To bridge this gap, we propose the Reflective Experience-Augmented Tutoring (REAT) framework, which couples experience distillation from historical dialogues with real-time adaptive retrieval. Driven by a multi-agent Observer-Critic-Mentor (OCM) distillation pipeline, REAT reviews past conversational trajectories and distills raw interactions into structured, problem-agnostic pedagogical experiences. During live tutoring, a state-aware retrieval module injects these curated experiences to provide adaptive scaffolding based on the student's cognitive state. Experiments demonstrate that the proposed framework significantly outperforms both prompt-only and supervised fine-tuning (SFT) baselines, particularly in improving complex, low-scoring tutoring scenarios. Crucially, the distilled experiences exhibit robust generalization across diverse model architectures and mathematical datasets.

↑