arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用大语言模型预测课程决策的后果

Anticipating the Consequences of Curriculum Decisions with Large Language Models

Octavio Pappalardo, Nathan Herr, Tim Rocktäschel

arXiv 2610.04604首次发表:更新:

发表机构

University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将自动课程学习建模为序贯决策问题,提出结合大语言模型信息与在线学习进度估计的方法,在Craftax基准上验证了其能更好地预测课程决策后果,尤其对单个目标任务优化效果显著。

AI 中文摘要

自动课程学习可以通过随时间选择呈现给智能体的训练经验来提高强化学习的有效性。然而,预测此类决策的后果可能很困难。我们将自动课程学习分析为一个序贯决策问题,突显了决定课程决策价值的数量与通常用于指导这些决策的局部学习信号所捕获的信息之间的差距。接着,我们研究大语言模型(LLMs)是否能够利用关于学习问题的更丰富信息来更好地预测课程决策的后果。我们提出了一种方法,将在线学习进度估计与LLM提供的估计相结合,这些估计包括:(i)在每个任务上学习的潜在下游收益,以及(ii)当前直接训练某个任务是否可能产生进展。我们在Craftax中针对不同课程目标,对包含256个文本目标的定制基准上评估了该方法。我们观察到,在优化单个目标任务时收益最强。当在整个任务集上优化时,收益因具有不同跨任务迁移机制的智能体而异,从学习速度的适度提升到持续到训练结束的更大收益不等。

英文摘要

Automatic curriculum learning can improve the effectiveness of reinforcement learning by selecting the training experiences presented to the agent over time. Predicting the consequences of such decisions can, however, be difficult. We analyze automatic curriculum learning as a sequential decision-making problem, highlighting a gap between the quantities that determine the value of curriculum decisions and the information captured by local learning signals commonly used to guide them. We then investigate whether Large Language Models (LLMs) can exploit richer information about the learning problem to better anticipate the consequences of curriculum decisions. We introduce a method that combines online learning-progress estimates with LLM-informed estimates of (i) the potential downstream benefits of learning on each task and (ii) whether direct training on a task is currently likely to produce progress. We evaluate the approach on a custom benchmark of 256 textual goals in Craftax under different curriculum objectives. We observe the strongest gains when optimizing for individual target tasks. When optimizing across the full task set, the benefits vary across learners with different mechanisms for cross-task transfer, ranging from modest improvements in learning speed to larger gains that persist through the end of training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑