发表机构
Sungkyunkwan University(成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EVOL利用知识追踪模拟器通过进化搜索合成专家示范,并训练免部署策略,在多个数据集上超越8种基线,解决学习路径推荐中的稀疏奖励问题。
AI 中文摘要
强化学习(RL)在学习路径推荐(LPR)中面临两个相互关联的障碍。首先,策略必须在不获得中间反馈的情况下承诺一个包含L个概念的序列,这产生了一个随L超指数增长且仅在最后一步提供奖励的组合搜索空间。其次,专家学习路径本是稀疏奖励强化学习的自然解决方案,但在教育数据中并不存在,因为学生日志记录的是学习者实际做了什么,而非他们本应做什么。我们通过引入机器人学中基于模拟器的示范学习方法来应对这两个障碍:知识追踪模拟器既用于通过进化搜索合成每个学习者的专家示范,也用于训练一个免部署策略,该策略将这些示范提炼为一个前馈学习者。我们的框架EVOL通过一个不对称的演员-评论家结构实例化了这一流程,其中演员承诺进行符合部署现实的盲目规划,而评论家在训练期间利用特权的模拟器状态。在三个数据集(ASSIST15、Junyi和EdNet;39-189个概念)以及路径长度L=5、10和20上,EVOL超越了涵盖启发式、顺序、强化学习、图增强强化学习和LLM增强方法的8个基线。我们进一步比较了三种模仿策略(BC、AWR和DAPG),并表明最终性能由进化专家的质量决定,而非特定的模仿目标。
英文摘要
Reinforcement learning (RL) for learning path recommendation (LPR) faces two coupled obstacles. First, the policy must commit to a sequence of L concepts without intermediate feedback, producing a combinatorial search space that grows super-exponentially with L and provides reward only at the final step. Second, expert learning paths would be the natural cure for sparse-reward RL, but they do not exist in educational data, because student logs record what learners did, not what they should have done. We address both obstacles by importing a recipe from simulator-based demonstration learning in robotics: the knowledge tracing simulator is used both to synthesize per-learner expert demonstrations through evolutionary search and to train a deployment-free policy that distills these demonstrations into a feed-forward learner. Our framework, EVOL, instantiates this pipeline with an asymmetric actor-critic where the actor commits to deployment-realistic blind planning while the critic exploits the privileged simulator state during training. Across three datasets (ASSIST15, Junyi, and EdNet; 39-189 concepts) and path lengths L = 5, 10, and 20, EVOL surpasses 8 baselines spanning heuristic, sequential, RL, graph-enhanced RL, and LLM-enhanced methods. We further compare three imitation strategies (BC, AWR, and DAPG) and show that final performance is governed by the quality of evolutionary experts rather than by the particular imitation objective.
CommentsAccepted at the 35th ACM International Conference on Information and Knowledge Management (CIKM 2026)