发表机构
Amazon Advertising Foundations; Amazon Web Services Agentic AI(亚马逊广告基金会; 亚马逊网络服务代理人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究企业编码代理将自然语言分析请求转换为可执行代码时的问题,提出两阶段训练后方法 CRAFT,先进行模式剥离的 PLAN 监督微调,再用执行形状的强化学习,评估显示其在多方面有提升并减少负担。
AI 中文摘要
企业编码代理将自然语言分析请求转换为通过专有 API、模式和度量定义的可执行代码。然而,当前将详尽的模式和工具文档注入每个提示的部署模式会增加推理开销,使模式演化复杂化,并破坏多轮分析中的可靠性。我们研究是否可以通过训练后获取稳定的模式知识和工具使用行为,同时保持面向生产分析所需的一致性。我们提出了 CRAFT,这是一种用于基于模式的编码代理的两阶段训练后方法。首先,模式剥离的 PLAN 监督微调从经过验证的轨迹中学习域结构计划和可执行行为,而无需详尽的提示时模式注入。其次,执行形状的强化学习调整工具选择、代码质量、计划代码一致性以及从失败执行中恢复的策略。训练轨迹通过结合执行验证的数据完整性检查和 LLM 法官推理审计的三门过滤器进行策划。我们评估 CRAFT 在广告分析中的计划推出,涵盖活动绩效分析、度量深入分析、实体级绩效分析和多轮分析细化。企业评估环境将测试版 API 作为面向代理的工具表面,涵盖 25 个模式链接的核心实体和 30 个代理工作流程。相对于填充模式的基线,CRAFT 将综合代理分数提高了 +9.6 个百分点,一致性提高了 +4.1 个百分点,多轮连贯性提高了 +4.2 个百分点,同时将输入令牌负担减少了约 9 倍,模式发现循环减少了多达 5 倍。我们还报告了企业环境中多轮工具使用强化学习所需的部署权衡、奖励塑造限制和训练基础设施扩展。
英文摘要
Enterprise coding agents translate natural-language analytical requests into executable code over proprietary APIs, schemas, and metric definitions. Yet the prevailing deployment pattern injecting exhaustive schema and tool documentation into each prompt increases inference overhead, complicates schema evolution, and undermines reliability in multi-turn analysis. We investigate whether stable schema knowledge and tool-use behavior can instead be acquired through post-training while preserving the consistency required for production-facing analytics. We present CRAFT, a two-stage post-training recipe for schema-grounded coding agents. First, schema-stripped PLAN supervised fine-tuning learns domain-structured plans and executable behaviors from validated trajectories without exhaustive prompt-time schema injection. Second, execution-shaped reinforcement learning aligns the policy for tool selection, code quality, plan-code consistency, and recovery from failed executions. Training trajectories are curated through a Tri-Gate filter combining execution validation, data-integrity checks, and LLM-judge reasoning audit. We evaluate CRAFT for planned rollout in advertising analytics, covering campaign performance analysis, metric drill-downs, entity-level performance analysis, and multi-turn analytical refinement. The enterprise evaluation environment incorporates beta APIs as the agent-facing tool surface and spans 25 schema-linked core entities and 30 agentic workflows. Relative to a schema-stuffed baseline, CRAFT improves composite Agent Score by +9.6 pp, consistency by +4.1 pp, and multi-turn coherence by +4.2 pp, while reducing input-token burden by approximately 9x and schema-discovery loops by up to 5x. We further report deployment tradeoffs, reward-shaping limitations, and training-infrastructure extensions required for multi-turn tool-use reinforcement learning in enterprise settings.