arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Palmyra x6技术报告:一种通过锚定监督微调进行后训练的具工具使用能力的智能体模型

Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning

Peng Du, Kiran Kamble, Rakshith Vasudev, Zhizhuo Yang, Rohith Nadimpally, Arjun Krishna, Waseem Alshikh, Daniel M. Bikel

arXiv 2608.16620首次发表:更新:

发表机构

Writer AI Research, Writer, Inc.(Writer AI研究院,Writer公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Palmyra x6是一款针对企业级智能体任务优化的大语言模型,通过锚定监督微调后训练构建,在公开基准测试及偏见安全评估中表现优异。

AI 中文摘要

Palmyra x6是一款针对企业级智能体任务优化的大语言模型。该模型通过在经验证的合成工具使用轨迹的紧凑语料库上,采用Muon与Adam混合优化器进行锚定监督微调,对混合专家基础模型进行后训练构建而成。该训练方案刻意保守且可控:共使用626条轨迹,训练1个epoch,采用低学习率,并对冻结的基础模型施加KL锚定约束。该模型相比Writer Agent的先前默认模型取得了显著提升,在公开基准测试中与多款近期模型表现相当,在BFCL Core上得分0.785,为最高值,且在同组六个基准测试中取得最高平均得分。此外,该模型在偏见与安全评估中表现出与对比模型相当或领先的水平。

英文摘要

Palmyra x6 is a large language model optimized for use with enterprise-oriented agentic tasks. The model was built by post-training a Mixture-of-Experts base model with Anchored Supervised Fine-Tuning on a compact corpus of verified, synthetic tool-use trajectories, optimized with a Muon + Adam hybrid. The recipe is deliberately conservative and deliberately controlled: 626 trajectories, a single epoch, a low learning rate, and a KL anchor to the frozen base. The model shows substantial gains over the previous default model for Writer Agent, and compares favorably with several recent models on public benchmarks, scoring the highest on BFCL Core at $0.785$ and posts the highest six-benchmark mean of the cohort. Furthermore, the model has shown itself to be competitive or leading relative to comparators in our bias and safety evaluations.

Comments12 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑