DENSE:将智能体轨迹蒸馏为基于证据的捷径树以进行自我精炼
DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
- Fudan University(复旦大学)
- Meituan Longcat Team(美团龙猫团队)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
DENSE将智能体执行轨迹蒸馏为基于证据的嵌套捷径树,在无结果标签下实现自我精炼,在Terminal-Bench 2.1上显著提升严格通过率并减少token消耗。
AI中文摘要:
在线智能体部署会产生大量的执行轨迹,而针对特定任务的验证和专家标注成本高昂且难以扩展。我们研究如何在没有事后结果标签的情况下,将这些轨迹蒸馏为可复用的反馈,利用其中关于局部进展、恢复和未完成需求的信息。我们提出了DENSE(从嵌套子任务执行中蒸馏证据),该方法将这些证据组织为基于证据的嵌套捷径树。DENSE压缩冗余尝试,利用恢复证据协调跨层级的问题,并总结已完成的分支同时扩展未解决的分支,将可复用的进展与剩余义务联系起来。我们引入了REFIT,一种源配对协议,在事后结果盲视下比较来自共享初始轨迹的反馈,并重置环境和模型上下文以对相同任务进行新的尝试。在Terminal-Bench 2.1上,DENSE在四个接收模型中取得了测试的非特权反馈方法中最高的严格通过率。相对于初始执行,严格通过率提高了7.12-15.64个百分点,重跑中观察到的接收者token减少了19.0-43.6%。GPT-5.5消融实验支持将嵌套子任务分析与捷径构建和问题协调相结合。这些发现指向通过基于证据的轨迹复用实现智能体自我精炼,减少对外部监督的依赖。
英文摘要:
Online agent deployments accumulate execution trajectories at massive scale and behavioral diversity, for which predefined annotation criteria hardly exist. Extracting useful evidence therefore demands costly manual annotation or verifier signals that fail to scale, leaving valuable evidence buried among redundant, incomplete, and failed executions. This raises a question: without post-execution rewards or correctness labels, how can reusable experience be distilled from the trajectories themselves? To address this challenge, we introduce DENSE (Distilling Evidence from Nested Subtask Executions), which organizes trajectory-derived evidence into nested shortcut trees. By consolidating redundant attempts, identifying resolved subtasks, and retaining useful steps alongside outstanding requirements, DENSE transforms noisy execution traces into structured and reusable task-solving feedback. To evaluate whether such feedback helps agents retry the same task, we design REFIT, which measures success-rate changes between the initial attempt and feedback-guided retries. Among feedback methods without external outcome supervision, DENSE achieves the highest strict pass rate across four agent models on Terminal-Bench 2.1, improving over initial attempts by 7.12-21.81 percentage points with 19.0-43.6% fewer agent tokens on retries. In addition, on hard tasks DENSE consistently outperforms self-reflection in cumulative pass rate across multiple feedback iterations on all four models, demonstrating its strong potential for continual agent self-improvement.