arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

经验时代中的自我发现强化学习:学习历史是资产还是负担?

Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?

Haomin Luo

arXiv 2609.35897首次发表:更新:

发表机构

University of Cambridge(剑桥大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次对自我发现的强化学习规则进行因果机制审计,揭示循环历史在扩展奖励尺度、解耦历史内容与维护、以及环境变化下控制重放保留三方面的影响,为自进化RL算法奠定审计标准。

AI 中文摘要

追求通用智能的递归自我改进(RSI)在宏观层面的语言模型扩展与“经验时代”的交互驱动原则之间产生了分歧。然而,任何自我改进架构最终都依赖于其底层优化引擎:如果通用智能需要从基于交互的学习中获取,那么强化学习(RL)更新规则本身必须具备累积适应的能力。尽管算法自我发现已产生超越PPO并达到SOTA基准性能的Disco103,但其内部更新机制仍是一个未被审视的黑箱。我们首次对自我发现的RL规则进行了因果机制审计,直接围绕经验时代的五大支柱构建:扩展视野、基于grounded的奖励尺度、持续流、生命周期内变化以及探索深度。通过在外科手术式地固定、冻结和移植循环状态的同时保持元参数不变,我们测试了学习历史何时作为资产或负担。三项发现组织了本次审计:(1)循环历史积极扩展了可用的奖励尺度,在零固定下维持了六十年窗口,而固定下仅三年。(2)将历史内容与其维护解耦揭示,不匹配历史的惩罚源于持续钳制;允许导入状态自然演化可减轻此负担。(3)在环境变化下,控制重放保留逆转了相对于DQN的表面适应优势,表明外部数据周转可能混淆内部可塑性。通过能力阈值验证并移植到第二个规则(OPEN),这项工作将宏观RSI雄心扎根于微观学习动态,为下一代自进化RL算法建立了基础审计标准。

英文摘要

The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-driven principles of "Era of Experience". Yet, any self-improving architecture ultimately rests upon its underlying optimization engine: if general intelligence requires learning from grounded interaction, the reinforcement learning (RL) update rule itself must be capable of cumulative adaptation. While algorithm self-discovery has produced Disco103 that surpassed PPO to achieve SOTA benchmark performance -- its internal update machinery remains an uninspected black box. We present the first causal mechanistic audit of a self-discovered RL rule, structured directly around the five pillars of the Era of Experience: extended horizon, grounded reward scales, continuing streams, within-lifetime change, and exploration depth. By surgically pinning, freezing, and transplanting recurrent states while holding meta-parameters fixed, we test when learning history acts as an asset or a burden. Three findings organize the audit: (1) Recurrent history actively expands usable reward scales, sustaining a six-decade window versus three under zero-pinning. (2) Decoupling historical content from its maintenance reveals that the penalty of mismatched history stems from perpetual clamping; allowing imported state to evolve naturally attenuates this burden. (3) Under environmental change, controlling replay retention reverses the apparent adaptation advantage over DQN, demonstrating that external data turnover can confound internal plasticity. Validated through capability thresholds and ported to a second rule (OPEN), this work grounds macro-RSI ambitions in micro-level learning dynamics, establishing a foundational audit standard for next-generation, self-evolving RL algorithms.

Comments52 pages, 17 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑