发表机构
Tongji University; Shanghai Artificial Intelligence Laboratory; Shanghai Jiaotong University; Peking University; Nanjing University(同济大学; 上海人工智能实验室; 上海交通大学; 北京大学; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SERA将评估能力内化为递归智能体自身,通过训练生成对齐的评分细则来提升分解与求解能力,在TextCraft-Synth和TextWorld-Sync上分别平均提升5.38和13.14分,推理时树搜索再增2.43分。
AI 中文摘要
递归语言模型智能体将任务分解,并将子任务委托给同一策略的子实例,从而形成一棵工作树。然而,训练它们很困难:最终结果可验证,但自行发明的中间子任务数量众多且没有真实标签。现有方法使用验证器或评判器对每个节点打分,这在规模上成本高昂,且对分解质量不敏感。我们认为,递归智能体必须在一组权重内学习三种耦合能力:将问题分解为子任务、解决子任务以及评估结果,每种能力都需要各自的训练信号。SERA(自评估递归智能体)将评估转变为策略自身的一种学习能力。在委派之前,父智能体为每个子任务编写一份加权成功标准评分细则;然后,针对可验证结果的排序目标训练评分细则生成,使标准能够追踪真实的子任务成功。此外,一种互补的叶覆盖信号为任务分解提供直接信用。我们的核心发现是,训练策略生成对齐的评分细则是带来收益的关键:因为相同的权重既用于评估又用于执行,学习评判子任务会增强智能体解决它们的能力。值得注意的是,外部监督也减少了:评判器仅用于训练评分细则生成器,而求解则针对智能体自身的评分细则分数进行训练,这优于直接使用评判器。在训练之外,学习到的评分细则还可作为推理时树搜索的选择器。在TextCraft-Synth和TextWorld-Sync上,SERA相比强递归智能体基线平均提高了5.38分和13.14分,推理时的评分细则引导树搜索在TextWorld-Sync上额外增加了2.43分。
英文摘要
Recursive language-model agents decompose tasks and delegate subtasks to child instances of the same policy, forming a tree of work. Training them, however, is hard: the final outcome is verifiable, but the self-invented intermediate subtasks are numerous and carry no ground truth. Existing methods score each node with a verifier or judge, which is costly at scale and blind to decomposition quality. We argue that a recursive agent must learn three coupled capabilities within one set of weights: decomposing problems into subtasks, solving them, and evaluating the outcomes, each requiring its own training signal. SERA (Self-Evaluating Recursive Agents) turns evaluation into a learned capability of the policy itself. Before delegating, the parent writes a rubric of weighted success criteria for each child subtask; a ranking objective against verified outcomes then trains rubric generation so that the criteria track genuine subtask success. In addition, a complementary leaf-coverage signal provides direct credit for task decomposition. Our central finding is that \emph{training} the policy to generate aligned rubrics is what drives the gains: because the same weights both evaluate and execute, learning to judge subtasks sharpens the agent's ability to solve them. Notably, external supervision is also reduced: the judge is consulted only to train the rubric generator, while solving is trained against the agent's own rubric scores, which outperform direct use of the judge. Beyond training, the learned rubric doubles as an inference-time selector for tree search. On TextCraft-Synth and TextWorld-Sync, SERA improves over strong recursive-agent baselines by 5.38 and 13.14 points on average, and rubric-guided tree search at inference adds a further 2.43 points on TextWorld-Sync.