发表机构
Shanghai Jiao Tong University; Tencent Youtu Lab(上海交通大学; 腾讯优图实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对开放式长文本生成任务轨迹数据稀缺的问题,提出RetroGen框架,利用预训练数据中丰富的高质量最终产物重构轨迹,训练模型以提升证据支撑等任务表现。
AI 中文摘要
轨迹数据对于训练大型语言模型以提升智能体能力正变得愈发重要。与编码或数学等可验证领域不同,为开放式任务扩展轨迹数据难度大得多,因为这类任务缺乏单一的真实值,且标注或验证成本高昂。本文提出RetroGen,一种回溯过程监督的自改进框架。我们的核心观察是:尽管专家轨迹稀缺,但高质量的最终产物(如文献综述、分析师报告和法律判决)在预训练数据中极为丰富,可视为生成这些产物的证据搜寻过程的压缩痕迹。RetroGen从专家产物中重构候选潜在轨迹,同时依据产物和支撑证据对其进行验证,并利用自身成功的重构数据训练模型,无需更强模型提供的轨迹数据。实验表明,RetroGen可提升证据支撑、忠实合成及长文本证据搜寻智能体任务的表现。
英文摘要
Trajectory data is getting more vital for training large language models for boosting the agentic abilities. Unlike the verifiable domains such as coding or mathematics, scaling trajectory data for open-ended tasks is much more difficult because these tasks lack singular ground truth and are costly to annotate or verify. In this paper, we propose RetroGen, a self-improving framework of retrospective process supervision. Our key observation is that although expert trajectories are scarce, high-quality final artifacts such as literature reviews, analyst reports and legal judgments, are abundant in pre-training data and can be viewed as compressed traces of the evidence-seeking processes that produced them. RetroGen reconstructs candidate latent trajectories from expert artifacts, verifies them against both the artifact and supporting evidence, and trains models on their own successful reconstruction data, without requiring trajectory data from stronger models. Experiments show that RetroGen improves grounding, faithful synthesis, and long-form evidence-seeking agent tasks.