发表机构
Massachusetts Institute of Technology; Hong Kong Polytechnic University; ByteDance Inc.; National University of Singapore; Stanford University; University of Washington; University of California, Berkeley; Tsinghua University; Singapore-MIT Alliance for Research and Technology(麻省理工学院; 香港理工大学; 字节跳动公司; 新加坡国立大学; 斯坦福大学; 华盛顿大学; 加州大学伯克利分校; 清华大学; 新加坡-麻省理工学院研究与技术联盟)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Guidance-TTT方法,通过测试时训练紧凑指导模型提出高层策略变更,冻结执行模型实现,在多个领域超越先前最佳方案。
AI 中文摘要
开放式科学发现通常需要反复提出和评估候选解决方案。基于LLM的系统可以通过从验证器反馈中生成和完善可执行解决方案来支持这一过程。诸如TTT-Discover之类的方法使用测试时训练(TTT)从验证器反馈中更新生成解决方案的LLM,调整其生成策略以改进针对目标问题的后续提案。然而,当可靠执行需要大型模型时,这变得昂贵,因为训练必须在重复生成长而结构化输出的同时维护梯度、优化器状态和策略统计。这也使信用分配复杂化:结果级验证器反馈必须共同评估高层策略及其低层实现。在这项工作中,我们引入了Guidance-TTT,它分离了这些角色。一个紧凑的指导模型在测试时被训练以提出高层策略变更,而一个冻结的执行模型将其实现为完整的可执行解决方案。在每一步中,系统选择一个有希望的先前发现的解决方案,提出变更,执行并验证它,并仅使用自适应组相对RL目标更新指导模型。这将测试时学习集中在短策略决策上,同时保留了一个明显更强模型的实现能力而无需适应它。在没有网络访问的情况下,Guidance-TTT在四个不同领域产生了强解决方案:组合优化(Polyomino Packing)、启发式编程(AHC058)、机器学习(Lasso)和GPU内核优化(TriMul)。在这些任务中,它优于先前工作中报告的最佳解决方案,同时与公共在线排行榜上的最先进结果保持竞争力。代码可在https://this https URL获取。
英文摘要
Open-ended scientific discovery often requires repeatedly proposing and evaluating candidate solutions. LLM-based systems can support this process by generating and refining executable solutions from verifier feedback. Methods such as TTT-Discover use test-time training (TTT) to update the solution-generating LLM from verifier feedback, adapting its generation policy to improve subsequent proposals on the target problem. However, this becomes expensive when reliable execution requires a large model, since training must maintain gradients, optimizer states, and policy statistics while repeatedly generating long, structured outputs. It also complicates credit assignment: outcome-level verifier feedback must jointly evaluate the high-level strategy and its low-level implementation. In this work, we introduce Guidance-TTT, which separates these roles. A compact guidance model is trained at test time to propose high-level strategic changes, while a frozen execution model implements them as complete executable solutions. At each step, the system selects a promising previously discovered solution, proposes a change, executes and verifies it, and updates only the guidance model using an adaptive group-relative RL objective. This concentrates test-time learning on short strategic decisions while retaining the implementation capability of a substantially stronger model without adapting it. Without web access, Guidance-TTT produces strong solutions across four distinct domains: combinatorial optimization (Polyomino Packing), heuristic programming (AHC058), machine learning (Lasso), and GPU kernel optimization (TriMul). Across these tasks, it outperforms the best solutions reported in prior work while remaining competitive with state-of-the-art results on public online leaderboards. Code is available at https://github.com/Human-Agent-Society/reef/tree/guidance-ttt-support.
Comments35 pages, including references and appendice