arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Agentic-TTT:为测试时训练训练测试时策略

Agentic-TTT: Training test-time policy for test-time training

Jiahao Lu, Mohan Kankanhalli

arXiv 2610.12002首次发表:更新:

发表机构

NUS AI Institute; National University of Singapore(新加坡国立大学AI研究院; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Agentic-TTT学习测试时策略管控TTT决策,可提升模型效用、权衡计算资源并泛化到新领域,推动模型自主自我改进。

AI 中文摘要

测试时训练(Test-time training, TTT)利用测试输入衍生的信号调整大语言模型(LLM)的参数,在国际数学奥林匹克(IMO)竞赛或指定开放问题等预设场景中可实现显著性能提升。TTT将部署阶段的经验转化为参数更新,为模型层面的自我改进提供了直接机制。然而TTT并非在所有场景下都适用:不同的TTT算法仅适配特定场景,若应用不匹配的方法,可能会浪费测试时计算资源,甚至损害模型性能。因此,这种参数层面的自我改进需要具备智能体能力:模型必须判断何时适合采用TTT、调用哪种算法,以及是否可复用现有技能。为填补这一空白,我们提出Agentic-TTT,该方法学习一种测试时策略来管控上述决策。Agentic-TTT将TTT流程转化为可调用工具,把积累的技能视为不断演进的部署环境,并利用其决策产生的观测效用增益来训练策略。在我们的基准测试中,Agentic-TTT的效用较骨干模型提升近一倍,学会了在效用与计算资源间进行权衡,且能泛化到训练过程中未见过的领域。这些结果共同指向自主自我改进的方向:即能够决定如何从自身部署经验中学习的模型。

英文摘要

Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems. By turning deployment experience into parameter updates, TTT provides a direct mechanism for model-level self-improvement. Yet TTT is not universally beneficial: each TTT algorithm works in different settings, and applying an ill-suited method could waste test-time compute or even damage model performance. Therefore, such parameter-level self-improvement requires agency: the model must decide when TTT is warranted, which algorithm to invoke, and whether an existing skill can be reused. To fill this gap, we introduce Agentic-TTT, which learns a test-time policy to govern those decisions. Agentic-TTT turns TTT procedures into callable tools, treats accumulated skills as an evolving deployment environment, and trains its policy using the observed utility gains from its decisions. On our benchmark, Agentic-TTT nearly doubles the utility over the backbone model, learns to trade off utility against compute, and generalizes to domains unseen during training. Together, these results point toward autonomous self-improvement: models that can decide how to learn from their own deployment experience.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑