arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23911cs.AIcs.LG

PROOF-Gen:从优化数据到更好的蒸馏

PROOF-Gen: From Optimized Data to Better Distillation

Anh Ta, Junjie Zhu, Shahin Shayandeh

首次发表
浏览论文内容

中文总结 AI 辅助

PROOF-Gen通过每场景提示优化从教师工具调用失败中恢复93%的错误轨迹,微调后提升了Qwen3-4B等模型的工具调用性能及部署流程的轨迹质量。

中文摘要 AI 辅助

对教师生成的轨迹进行监督微调,是将工具调用能力蒸馏为可部署模型的标准第一阶段。驱动已部署工具调用智能体的后训练流程,每日或每周重复此阶段,每个周期都需支付前沿教师模型的成本,而该流程采用的是生成-过滤机制(保留教师的合格轨迹,丢弃其余轨迹),且每个周期都会留下相同的困难场景,因为失败无法提供信号。在τ2-bench上,57%的教师试验失败,其中三分之二是差一点成功的情况(多数工具调用正确,仅因一个关键错误导致失败)。我们提出PROOF-Gen(Per-scenario Reflective Optimization to Overcome Failed Generation,用于克服失败生成的每场景反射优化),它通过每场景提示优化从这些失败中恢复正确轨迹。对于每个失败的任务,反射器分析执行轨迹和评估反馈,然后编写纠正指导,引导教师生成合格的轨迹。指导内容在训练前被剥离,因此学生从干净的演示中学习,无特定任务的支架。在τ2-bench上,每场景优化恢复了93%的失败场景。在合并数据上微调后,Qwen3-4B-Instruct-2507的Pass^1从0.132提升至0.529,Gemma 4 E4B-it在BFCL v4多轮任务上获得7.2个百分点的提升。在已部署流程中,该方法将轨迹质量的目标完成率提升6.3个百分点,并可迁移到已部署的设备端模型(目标完成率提升1.5个百分点;响应质量指标提升1.7至5.0个百分点),且在所有语言环境中均有正向迁移(非英语环境平均提升1.48个百分点)。

英文摘要

Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run this stage on a daily or weekly cadence, paying the frontier-teacher cost each cycle, yet the mechanism is generate-and-filter (keep the teacher's passing trajectories, discard the rest) and each cycle leaves behind the same hard scenarios because failures supply no signal. On τ2-bench, 57% of teacher trials fail, two-thirds of them near-misses (most tool calls correct, undone by one decisive error). We introduce PROOF-Gen (Per-scenario Reflective Optimization to Overcome FailedGeneration), which recovers golden trajectories from these failures via per-scenario prompt optimization. For each failed task, a reflector analyzes the execution trace and evaluation feedback, then writes corrective guidance that steers the teacher to a passing trajectory. The guidance is stripped before training, so the student learns from clean demonstrations with no task-specific scaffold. On τ2-bench, per-scenario optimization recovers 93% of failed scenarios. Fine-tuned on the combined data, Qwen3-4B-Instruct-2507 improves from Pass^1=0.132 to 0.529 and Gemma 4 E4B-it gains +7.2pp on BFCL v4 multi-turn. In a deployed pipeline, the method lifts trajectory quality by +6.3pp goal completion and transfers to a deployed on-device model (+1.5pp goal completion; +1.7 to +5.0pp across response-quality metrics), with positive transfer in every locale (non-English average +1.48pp).

发表机构

  • Apple(苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

↑