驾驭学习实现可泛化的测试时自适应
Harness Learning Enables Generalizable Test-Time Adaptation
浏览论文内容
中文总结 AI 辅助
提出驾驭学习,通过强化学习训练提议者模型,利用执行反馈修订语言模型智能体的驾驭程序,实现无需参数更新的测试时自适应,并在推理与多跳问答任务上验证了其可泛化性。
中文摘要 AI 辅助
语言模型智能体由其模型和驾驭(harness)共同定义,驾驭是组织模型调用、工具使用和信息流的可执行程序。由于不同任务需要不同的方式来组织这些操作,驾驭需要利用当前任务的反馈进行自适应。我们引入了驾驭学习,它训练一个提议者模型,利用执行反馈来修改求解器的驾驭。我们将这一过程表述为对可执行程序的元学习,其中驾驭修订在基于梯度的自适应中扮演权重更新的角色。我们使用强化学习训练提议者,以修订后驾驭的任务性能作为奖励。在测试时,提议者利用在新任务上连续执行的反馈来优化驾驭,而不进行任何参数空间更新。在推理和多跳问答上的实验表明,驾驭学习提高了修订质量,并且测试时自适应的能力可迁移到未见过的任务。基于单个修订训练的策略能够在多轮中持续改进驾驭,而基于修订序列训练的好处因设置而异。这些发现指向了一条通往持续学习智能体的路径,该智能体将积累的经验转化为可泛化的改进。
英文摘要
A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver's harness using execution feedback. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, without performing any parameter-space update. Experiments on reasoning and multi-hop question answering show that harness learning improves revision quality and that the ability to adapt at test time transfers to unseen tasks. Policies trained on individual revisions can continue improving harnesses over multiple rounds, while the benefits of training on revision sequences vary across settings. These findings suggest a path towards continually learning agents that turn accumulated experience into generalizable improvements.