arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38446cs.LGcs.AIcs.CL

预训练与中期训练使什么可从奖励中学习?

What Pretraining and Midtraining Make Learnable from Rewards?

Chiwun Yang, Xiaoyu Li

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过理论分析和Qwen2.5实验,证明预训练提供信息、中期训练提供可执行计算,使奖励适应有效,并在多世界任务中显著提升成功率。

中文摘要 AI 辅助

奖励可以识别正确答案,但无法确定新输入所需的计算过程。我们研究预训练和中期训练如何提供信息和计算,使奖励适应变得有效。在顺序状态计算和上下文记忆中,我们刻画了那些在每一个训练奖励上一致但对保留答案要求不同的机制。任务无关的源观测解决了这种歧义。我们从指定的随机初始化出发,通过源预测和奖励适应在同一参数中构造有限采样的Adam路径,证明了预测如何获取执行或检索能力,以及奖励如何学习其任务特定的用途。使用预训练的Qwen2.5检查点的实验检验了这种分工。在八个世界中,使用正确源和第一操作监督训练的序列模型达到82.61%的成功率,而私有随机源对照仅为44.15%。记忆回放在奖励适应过程中保持检索能力,独立的八世界确认在匹配的替代检索训练后达到75.32%的任务成功率,而对照为49.86%。GSM8K和HotpotQA在奖励进入时的准确率、后续增益和最终性能上有所区分。总之,这些结果将信息获取、可执行计算和奖励引导的任务学习联系起来。

英文摘要

A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study how pretraining and midtraining supply the information and computation that make reward adaptation effective. In sequential state computation and contextual memory, we characterize mechanisms that agree on every training reward yet demand different held-out answers. Task-independent source observations resolve this ambiguity. We construct finite sampled Adam paths from specified random initializations through source prediction and reward adaptation in the same parameters, proving how prediction acquires execution or retrieval and rewards learn their task-specific use. Experiments with pretrained Qwen2.5 checkpoints test this division of labor. Across eight worlds, Sequential models trained with correct source and first-operation supervision reach 82.61% success, versus 44.15% for a private-random source control. Memory replay preserves retrieval during reward adaptation, and an independent eight-world confirmation achieves 75.32% task success versus 49.86% after matched alternative-retrieval training. GSM8K and HotpotQA separate accuracy at reward entry, subsequent gain and final performance. Together, these results connect information acquisition, executable computation and reward-guided task learning.

发表机构

  • City University of Hong Kong(香港城市大学)
  • University of New South Wales(新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑