arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Q学习实验室:通过学习者生成的轨迹分析教授强化学习

Q-Learning Lab: Teaching Reinforcement Learning Through Learner-Generated Trace Analysis

Ekkachai Jueng

arXiv 2607.10802首次发表:更新:

发表机构

Computer Science Program Faculty of Sciences and Liberal Arts Rajamangala University of Technology Isan(计算机科学系 法学院及文科院 瑞亚玛芒拉大学技术学院伊桑)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出Q学习实验室工具,通过学习者生成的轨迹分析教授表格Q学习。工具具可视化及记录功能,核心是学习-导出-分析循环。通过三项评估验证工具,还与现有工具比较、阐述教学法基础并提供教案,工具和代码公开。

AI 中文摘要

强化学习通常通过贝尔曼更新引入,但本科生往往觉得该方程抽象:他们看到策略箭头收敛,却很少观察每个值如何计算或为何选择某个动作。我们展示了Q学习实验室,这是一个无需安装的单文件、基于浏览器的双语(泰语/英语)工具,用于教授表格Q学习。除了常见的网格世界可视化,该工具还展示实时贝尔曼替换面板,记录每次转换并导出轨迹。核心贡献是学习-导出-分析循环,学习者运行代理、导出轨迹并自行分析。我们通过三项评估验证了该工具:与价值迭代真值对比学习值和策略的正确性;对超参数进行扫描;进行奖励编辑研究。我们还将该工具与现有网格世界可视化工具进行比较,描述其在实践教学法中的基础,并提供了一个50分钟的教案。该工具和所有实验代码均可公开获取。

英文摘要

Reinforcement learning is usually introduced through the Bellman update, yet the equation often remains abstract to undergraduates: they watch policy arrows converge but rarely observe how each value is computed or why an action is chosen. We present Q-Learning Lab, a single-file, browser-based, bilingual (Thai/English) tool for teaching tabular Q-learning that requires no installation. Beyond the usual gridworld visualization - color-coded Q-values and policy arrows on a $5 \times 5$ world - the tool exposes a live Bellman-substitution panel showing the numeric update at every step, and logs each transition, including the full pre-action Q-row, the greedy-versus-random decision under $\varepsilon$-greedy exploration, and wall-collision events, into an exportable trace. The central contribution is a learn-export-analyze loop: learners run their own agent, export the complete trace as CSV, and analyze it themselves, producing learning curves, value heatmaps, and visitation maps, turning a passive demonstration into a source of learner-generated data for reflective inquiry. We validate the tool without human-subject data through three complementary evaluations: (i) correctness of the learned values and policy against a value-iteration ground truth on the identical MDP; (ii) hyperparameter sweeps over $α$, $γ$, and $\varepsilon$ showing that every pedagogical claim the tool makes is reproducible; and (iii) a reward-editing study that uses the ground-truth optimal policy to separate two behaviorally identical but diagnostically opposite failure modes - an exploration failure versus genuine reward misspecification - that a single edited reward can produce. We also compare the tool against existing gridworld visualizers, describe its grounding in learning-by-doing pedagogy, and include a 50-minute lesson plan. The tool and all experiment code are openly available.

Comments12 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑