arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FLARE:一种通过生成式奖励模型实现长时程编码智能体全生命周期密集监督的范式

FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model

Jingxuan Xu, Gang Wu, Yanan Wu, Yutao Mou, Songwei Yu, Tianzhuang He, Zhengshuo Gong, Zhao Liu, Zihang Xu, Wenqiang Zhu, Xinping Lei, Weihao Li, Yuhui Bai, Zhongqiu Wang, Yan Wu, Ariel Deng

arXiv 2609.23808首次发表:更新:

AI 中文总结

FLARE提出一种基于生成式奖励模型的全生命周期密集监督范式,通过因果诊断和实时风险反馈,在推理与训练中优化长时程编码智能体,显著降低计算开销并提升性能。

AI 中文摘要

虽然测试时扩展增强了大型语言模型(LLM)智能体在长时程软件工程(SWE)中的表现,但稀疏的二元奖励(通过/失败)造成了严重的信用分配危机,并浪费了失败的探索轨迹。当前的轨迹优化和扩展方法成本高昂且结构受限,依赖于无因果诊断的启发式状态重用,或缺乏可操作在线指导的延迟标量评分。我们提出了FLARE(全生命周期对齐与奖励引擎),一种由轻量级生成式奖励模型(GRM)驱动的新型密集监督范式。首先,RADAR,一个离线因果感知诊断框架,通过因果链回溯提取高保真、无后见之明的监督信号,以蒸馏出一个提供实时、步骤级风险反馈的GRM。其次,FLARE利用该GRM在智能体的整个生命周期中持续优化它。在推理期间,FLARE充当主动脚手架,自主拦截高风险生成步骤以进行局部断点重新执行,大幅降低计算开销。在后训练期间,GRM的结构化信号作为监督微调(SFT)的过程监督重排序分数,以及强化学习(RL)的步骤级密集奖励,缓解了稀疏环境中的策略崩溃。大量评估表明,FLARE在智能体生命周期中建立了新的帕累托前沿:FLARE(N=1)在token消耗减少5倍的情况下优于全局展开(N=5)。将FLARE扩展到训练中,克服了长时程交互任务中的稀疏奖励问题,通过过程感知数据整理在SFT中实现了19.13%的相对性能提升,并在RL中实现了一致的9.19%的提升。

英文摘要

While test-time scaling enhances Large Language Model (LLM) agents in long-horizon software engineering (SWE), sparse binary rewards (Pass/Fail) create a severe credit assignment crisis and waste failed exploratory trajectories. Current trajectory optimization and scaling methods are costly and structurally limited, relying on heuristic state reuse without causal diagnosis or delayed scalar scoring without actionable online guidance. We propose FLARE (Full-Lifecycle Alignment and Reward Engine), a novel dense supervision paradigm driven by a lightweight Generative Reward Model (GRM). First, RADAR, an offline causal-aware diagnostic framework, extracts high-fidelity, hindsight-free supervision through causal-chain backtracking to distill a GRM providing real-time, step-level risk feedback. Second, FLARE uses this GRM to continuously optimize the agent across its entire lifecycle. During inference, FLARE acts as an Active Scaffold, autonomously intercepting high-risk generation steps for localized breakpoint re-execution, drastically reducing compute overhead. During post-training, the GRM's structured signals serve as process-supervised reranking scores for Supervised Fine-Tuning (SFT) and step-level dense rewards for Reinforcement Learning (RL), mitigating policy collapse in sparse environments. Extensive evaluations show that FLARE establishes a new Pareto frontier across the agent lifecycle: FLARE (N=1) outperforms Global Rollout (N=5) with a 5x reduction in token consumption. Extending FLARE to training overcomes the sparse reward problem in long-horizon interactive tasks, delivering relative performance gains of 19.13% in SFT through process-aware data curation and a consistent 9.19% improvement in RL.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑