训练你所部署的内容:面向生产编码的 token 忠实性后训练
Train What You Deploy:Token-Faithful Post-Training of a Production Coding
浏览论文内容
中文总结 AI 辅助
针对编码后训练的 token 与控制忠实性问题,提出感知忠实性的训练耦合框架及 C-DPPO 算法,在 Baize 系列模型上实现 3.0 点性能提升,验证了流程可靠性。
中文摘要 AI 辅助
现有编码与终端智能体的后训练流程存在严重的 token 与控制忠实性错误:简化的训练环境与生产部署不匹配,且从智能体日志进行的离线 token 重构会扭曲原始提示词,并将策略调用与背景模型操作混为一谈。我们提出一种感知忠实性的训练耦合框架,该框架保留训练器侧对原始提示词的采样,通过协商训练协议消除虚假模型调用,并将损失计算限制在具有闭失败保证的可验证 token 跨度内。我们进一步提出认证散度近端策略优化(Certified Divergence Proximal Policy Optimization,C-DPPO),其在标准 DPPO 的基础上建立了严格的双侧 TV 认证边界、自适应 K 规则、预算感知序列保证以及抗错误策略掩码。在匹配的 Baize5B 和 Baize10B 模型上,采用 TMax-100 上相同的训练与测试协议进行评估,C-DPPO 相较于标准 DPPO 在所有模型规模上均实现了一致的 3.0 点性能提升。认证审计验证了我们的认证训练流程的可靠性和完整操作覆盖范围。
英文摘要
Existing post-training pipelines for coding and terminal agents suffer severe token and control fidelity errors: simplified training environments mismatch production deployments, and offline token reconstruction from agent logs distorts original prompts and conflates policy calls with background model operations. We present a fidelity-aware training coupling framework that retains trainer-side sampling over original prompts, eliminates spurious model calls via a negotiated training protocol, and restricts loss computation to verifiable token spans with closed-failure guarantees. We further propose Certified Divergence Proximal Policy Optimization (C-DPPO), which establishes tight two-sided TV certification bounds, adaptive-K rules, budget-aware sequence guarantees, and error-robust policy masking atop standard DPPO. Evaluated on matched Baize5B and Baize10B models with identical training and test protocols on TMax-100, C-DPPO yields a consistent +3.0-point performance gain over standard DPPO across model scales. Certificate audits validate the reliability and full operational coverage of our certified training pipeline.
发表机构
- KunlunMeta(昆仑元)
机构由 AI 辅助整理,请以论文原文为准。