一点SFT,一点RL:强化学习何时帮助长周期广告智能体
A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents
浏览论文内容
中文总结 AI 辅助
本研究通过诊断教师支持与奖励余量,在长周期广告智能体中平衡SFT与RL,针对性RL提升7/8技能,最大增益+11.27点,并降低泄漏。
中文摘要 AI 辅助
企业分析智能体在分布式业务数据上解决长周期工具使用问题,需要检索、推理、API调用、代码执行以及对中间观测的适应。监督微调(SFT)校准工具语法和教师支持的行为,而强化学习(RL)可以探索超出演示的奖励支持行为;然而,如果统一应用,RL可能会扰动已经校准的技能。我们研究在模拟生产的beta API下如何平衡SFT和RL。我们观察到,在我们的受控实验中,检查点轨迹回顾性地分为三个机制:模仿(Imitation),其中SFT捕获了可靠的教师行为;提升(Lift),其中两个阶段都有帮助;以及发现(Discovery),其中有用的奖励可观测行为位于可靠的教师支持之外。我们前瞻性地利用这一点,使用教师支持和奖励可观测余量将特征路由到仅SFT、先SFT后RL、增加RL分配或进一步环境开发。在随后的18个特征特定实验中,该诊断预测了15/18个观察到的轨迹。在GPT-OSS 120B上,针对性的先SFT后RL相对于前沿控制(Control)在7/8个广告商技能上产生了正的点估计;五个正增益具有排除零的配对95%置信区间,而一个技能具有置信支持的回归。最大增益是非披露(+11.27点;95% CI [+9.72, +12.82])。一项单独的SME审计表明,相对于SFT,针对性的RL将标准泄漏从11.8%降低到2.9%,对抗性泄漏从22.9%降低到6.8%,同时保持可操作性(86.2%到85.7%)。在共享奖励和优化的匹配统一与针对性比较中,针对性的RL将七个技能的平均增量从+1.62提高到+3.57,同时使用减少43%的增量RL计算。
英文摘要
Enterprise analytics agents solve long-horizon tool-use problems over distributed business data, requiring retrieval, reasoning, API calls, code execution, and adaptation to intermediate observations. Supervised fine-tuning (SFT) calibrates tool syntax and teacher-supported behavior, whereas reinforcement learning (RL) can explore reward-supported behaviors beyond demonstrations; applied uniformly, however, RL can perturb already-calibrated skills. We study how to balance SFT and RL under production-mirroring beta APIs. We observe that, in our controlled experiment, checkpoint trajectories retrospectively separated into three regimes: Imitation, where SFT captured reliable teacher behavior; Lift, where both stages helped; and Discovery, where useful reward-observable behavior lay outside reliable teacher support. We leverage this prospectively, using teacher support and reward-observable headroom to route features to SFT only, SFT then RL, increased RL allocation, or further environment development. Across 18 subsequent feature-specific experiments, the diagnostic predicted 15/18 observed trajectories. On GPT-OSS 120B, targeted SFT then RL produced positive point estimates on 7/8 advertiser skills relative to a frontier Control; five positive gains had paired 95% confidence intervals excluding zero, while one skill had a confidence-supported regression. The largest gain was non-disclosure (+11.27 points; 95% CI [+9.72, +12.82]). A separate SME audit surfaced that targeted RL reduces standard leakage from 11.8% to 2.9% and adversarial leakage from 22.9% to 6.8% relative to SFT while preserving actionability (86.2% to 85.7%). In a matched uniform-versus-targeted comparison with shared rewards and optimization, targeted RL improved the seven-skill mean delta from +1.62 to +3.57 while using 43% less incremental RL compute.
发表机构
- Amazon Advertising(亚马逊广告)
- Amazon Web Services Agentic AI(亚马逊云科技智能体人工智能)
机构由 AI 辅助整理,请以论文原文为准。