arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越直接回答:通过启发式强化学习将教育大语言模型调整为苏格拉底式引导者

Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning

Xiaokun Wang, Siyu Song, Wentao Liu, Xiaodong Zou

arXiv 2607.22996首次发表:更新:

发表机构

East China Normal University; Shanghai Chuangjie Situo Information Technology Co., Ltd.; Shanghai Normal University(华东师范大学; 上海创杰思拓信息技术有限公司; 上海师范大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对教育场景中LLMs直接作答问题,提出HeuristicEdu两阶段管道,经监督热身和GRPO,利用特定对话及启发式奖励训练Qwen2.5 - 7B,提升了支架有效性,降低关键词泄漏,表明规模并非引发苏格拉底行为的唯一因素。

AI 中文摘要

在教育环境中部署的大语言模型(LLMs)通常表现为直接回答者,不像苏格拉底教学法规定的那样通过渐进式探究引导学生。我们提出HeuristicEdu,这是一个两阶段管道,通过监督热身和组相对策略优化(GRPO)使Qwen2.5 - 7B向苏格拉底辅导对齐。训练使用从直播平台重建的797个多轮中文儿童科学对话,有基于认知深度(R_cog)、好奇心参与度(R_eng)和直接性(R_dir)的启发式奖励以及对学生引入术语的K_query校正。引入支架有效性(SE)和对话深度(CD)来评估表面流利度之外的结果。在30个保留问题上,最佳GRPO变体将SE从30.0%提高到63.3%,并将关键词泄漏从30.0%降低到13.3%。未对齐的Qwen - 72B基线SE为0%,泄漏率为96.7%。

英文摘要

Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the opening turn instead of guiding students through progressive inquiry, as Socratic pedagogy prescribes. We present HeuristicEdu, a two-phase pipeline that aligns Qwen2.5-7B toward Socratic tutoring via supervised warm-up and Group Relative Policy Optimization (GRPO). Training uses SocraticEdu, 797 multi-turn Chinese children's science dialogues reconstructed from a live platform, with a heuristic reward over cognitive depth (R_cog), curiosity engagement (R_eng), and directness (R_dir), together with a K_query correction for student-introduced terms. We introduce Scaffolding Effectiveness (SE) and Conversation Depth (CD) to evaluate outcomes beyond surface fluency. On 30 held-out questions, the best GRPO variant improves SE from 30.0% to 63.3% and lowers keyword leakage from 30.0% to 13.3%. Notably, this best variant omits the directness penalty during optimization, suggesting that explicit anti-leakage terms can conflict with gradient-based behavioral alignment. An unaligned Qwen-72B baseline reaches 0% SE and 96.7% leakage, showing that scale alone does not induce Socratic behavior.

Comments12 pages, 2 figures, 5 tables; includes an appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑