arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Salesforce Koa:面向智能体工具使用的企业级语言模型

Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

Zixiang Chen, Sufeng Niu, Yingchi Liu, Wenting Zhao, Akshara Prabhakar, Shubham Mehrotra, Bin Bi, Zhujun Lan, Katherine Tan, Mohammad Ramezanali, Tulika Manoj Awalgaonkar, Monojit Banerjee, Jielin Qiu, Shiva Kumar Pentyala, Zhepeng Cen, Anupam Tripathi, Ali Ziaei, Regunathan Radhakrishnan, Darvish Lee Shadravan, Shelby Heinecke, Sitaram Asur, Jayesh Govindarajan, Silvio Savarese, James Zhu, Phil Mui, Huan Wang

arXiv 2609.15066首次发表:更新:

发表机构

Salesforce Agentforce & AI Research(赛富时 Agentforce 与人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Salesforce Koa通过规范驱动的强化学习后训练开放权重模型,提升多轮工具使用和智能体能力,在企业基准上超越专有基线。

AI 中文摘要

我们提出了Salesforce Koa,一个企业级语言模型,它通过对开放权重的Nemotron-3-Super-120B基础模型进行后训练,并使用组相对策略优化(GRPO)进行强化学习而构建。Salesforce Koa在公开数据和合成生成数据上训练,不涉及客户数据,旨在提升工具使用和智能体能力,同时保持强大的通用性能。其独特组成部分是一个模拟到奖励的流水线,该流水线将工作流规范扩展为基于人物角色的多轮任务,并为数据相关请求提供基于成功工具使用的任务解决奖励。对于企业领域,这些规范使用Agent Script编写,这是Salesforce用于构建Agentforce代理的声明式语言;对于公共工具使用领域,我们直接合成工作流结构。相同的模拟和基于奖励的机制驱动GRPO在两者上的应用。在公共工具使用、智能体推理和企业客户关系管理(CRM)基准测试中,Salesforce Koa相比其开放权重基础模型有所改进,在多轮工具使用上提升最为明显,并超过了强大的专有基线,但仍低于最强的前沿模型。这些结果表明,规范驱动的强化学习是将开放权重基础模型专业化用于企业智能体任务的实用途径。

英文摘要

We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO), and deployed in FP8 for production. Koa is trained only on public and synthetically generated data, and specialized for the agentic tool use that enterprise workflows demand: routing a request to the correct action, invoking the right tool with valid arguments, and completing multi-turn business tasks. The distinctive component of our pipeline is specification-driven task construction: declarative Agent Script specifications are expanded into persona-conditioned multi-turn environments whose rewards are grounded in successful tool use. Applied to enterprise CRM specifications, the same pipeline produces the in-domain training distribution on which Koa is specialized. On CRMAgentBench and the human-labeled production tool-calling set, Koa outperforms both its untuned open-weight base and GPT-4.1 and is competitive with the strongest frontier models. It reaches 87% Task Success Rate on CRMAgentBench (vs. GPT-4.1 at 82% and the base at 79%), is at or near the top of every metric on the human-labeled portion of an internal production benchmark, and preserves the base model's general capability on public benchmarks (Tau2Bench, BFCL). A controlled comparison with architecture and RL recipe held fixed shows that the additional in-domain RL stage improves argument accuracy and full tool-call success on the human-labeled enterprise benchmark.

Comments17 pages, 1 figure, 11 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑