发表机构
State University of New York at Buffalo(纽约州立大学布法罗分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DICE通过解耦干预决策与响应生成,利用干预价值指标和策略优化,在数学辅导中实现选择性干预,减少过度干预并提升效率。
AI 中文摘要
流畅的指导并不等同于有效的干预。大语言模型辅导系统通常被训练为生成下一个教师话语,这隐含地假设每个学生发言都需要回应。然而,我们的实验表明,这种做法混淆了辅导能力(说什么)与干预必要性(是否说)。我们提出了DICE框架,该框架通过首先选择明确的 pedagogical 动作,将干预决策与响应生成解耦。为了校准这一动作选择策略,我们定义了干预价值(IV),这是一种基于 rollout 的反事实指标,将每个动作与不干预进行比较。IV表明,许多规定的干预措施几乎没有或根本没有边际收益。我们进一步引入了DICE-Bench,一个多变体数学辅导基准,包含保留技能的问题变体,用于会话级评估。通过使用IV加权和KL正则化的策略优化,DICE学会了选择性干预,同时保持辅导效果。在模拟辅导会话中,DICE将过度干预率降至接近零,同时平均比现有的苏格拉底式辅导基线少用约3-4个回合引导学生得出正确答案。
英文摘要
Fluent guidance is not the same as useful intervention. LLM tutors are typically trained to generate the next teacher utterance, implicitly assuming that every student turn warrants a response. However, our experiments indicate that this conflates tutoring capability (what to say) with intervention necessity (whether to say it). We introduce DICE, a framework that decouples intervention decisions from response generation by first selecting an explicit pedagogical action. To calibrate this action selection policy, we define Intervention Value (IV), a rollout-grounded counterfactual metric that compares each action against non-intervention. IV shows that many prescribed interventions provide little or no marginal benefit. We further introduce DICE-Bench, a multi-variant math tutoring benchmark with skill-preserving problem variants for session-level evaluation. Using IV-weighted and KL-regularized policy optimization, DICE learns to intervene selectively while preserving tutoring effectiveness. In simulated tutoring sessions, DICE reduces the over-intervention rate to near zero while guiding students to correct solutions in approximately 3-4 fewer turns on average than existing Socratic tutoring baselines.
CommentsAccepted in NeurIPS 2026