发表机构
University of Illinois Urbana-Champaign; University of Maryland, College Park; Stanford University(伊利诺伊大学厄巴纳-香槟分校; 马里兰大学帕克分校; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CURIO提出好奇心驱动的测试时学习框架,用内在好奇心世界模型补充任务反馈,在数学发现和单细胞去噪任务上提升性能,支持好奇心作为探索信号。
AI 中文摘要
开放式发现需要从反复尝试中学习,同时继续探索其价值尚不明显的方向。使用冻结的大语言模型(LLM)进行搜索可以在上下文中重用先前的解决方案,但无法根据测试问题上的成功与失败来更新模型。强化学习(RL)能够实现这种适应;然而,强烈偏向高奖励轨迹可能会过早抑制低奖励但可能具有前景的方向。我们提出CURIO,一种好奇心驱动的测试时学习框架,用内在好奇心世界模型(ICWM)补充任务反馈。ICWM学习策略隐藏状态表示中的转移,并在策略的top-k选择之外的采样token上提供预测误差奖励。轮次归一化和退火权重调节它们对策略更新的贡献。在六个数学发现任务和单细胞去噪任务上,使用Qwen3骨干模型(从8B到235B),三次运行的平均值在五个数学目标上优于匹配的仅任务RL对照,在Circle Packing上匹配最佳报告性能,并在每个测试规模下在两个保留语料库上改善去噪Score和均方误差(MSE)。相对增益在Hadamard上达到18.3%,在去噪Score上达到10.8%。代码多样性测量显示生成程序之间的结构变化更大,支持好奇心作为开放式发现学习中互补的探索信号。
英文摘要
Open-ended discovery requires learning from repeated attempts while continuing to explore directions whose value is not yet apparent. Search with a frozen large language model (LLM) can reuse previous solutions in context, but cannot update the model from its successes and failures on the test problem. Reinforcement learning (RL) enables such adaptation; however, strongly favoring high-reward trajectories may suppress low-reward yet potentially promising directions too early. We introduce CURIO, a curiosity-driven test-time learning framework that complements task feedback with an Intrinsic Curiosity World Model (ICWM). The ICWM learns transitions in the policy's hidden-state representation and supplies prediction-error bonuses at sampled tokens outside the policy's top-k choices. Epoch normalization and an annealed weight regulate their contribution to the policy update. On six mathematical discovery tasks and single-cell denoising with Qwen3 backbones from 8B to 235B, three-run means improve over a matched task-only RL control on five mathematical objectives, match the best reported performance on Circle Packing, and improve denoising Score and mean squared error (MSE) on both held-out corpora at every tested scale. Relative gains reach 18.3% on Hadamard and 10.8% on denoising Score. Code-diversity measurements show greater structural variation among generated programs, supporting curiosity as a complementary exploration signal for learning in open-ended discovery.