当API“说错”语言:重新审视多语言工具使用的后训练
When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use
查看机构详情
- Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对多语言API调用中存在的参数语言不匹配问题,研究发现监督微调可实现接近或优于复杂强化学习方法的性能,强化学习仅能提供渐进式改进。
中文摘要 AI 辅助
大型语言模型(LLMs)在API调用中的可靠性在多语言场景下会下降。一种常见的失败情况是,模型选择了正确的工具,但生成的参数值语言不一致,我们将此称为参数语言不匹配(Argument Language Mismatch, ALM)。尽管这些输出在语义上是正确的,但在操作上是无效的,且未被标准API调用指标捕获。我们重新审视用于缓解ALM的后训练策略,在我们的基准测试中发现,监督微调(supervised fine-tuning, SFT)提供了很强的基线,大幅提高了参数语言一致性和端到端函数调用准确率。在模型选择一致的情况下,SFT的性能可与更复杂的强化学习(reinforcement learning, RL)方法相媲美,有时甚至超过后者。我们进一步研究带结构化、参数感知奖励的RL是否能带来额外益处。尽管诸如组相对策略优化(Group Relative Policy Optimization, GRPO)之类的方法可提高语言一致性并更好地保留通用推理能力,但这些收益是渐进式的,且在泛化和多目标权衡中最为显著。总体而言,我们的结果表明,多语言API定位的大部分性能可通过精心设计的监督训练实现,而RL仅能提供针对性而非根本性的改进。
英文摘要
The reliability of Large Language Models (LLMs) for API calling degrades in multilingual settings. A common failure occurs when a model selects the correct tool but generates argument values in an inconsistent language, which we term Argument Language Mismatch (ALM). Although semantically correct, such outputs are operationally invalid and not captured by standard API-calling metrics. We revisit post-training strategies for mitigating ALM and find that, in our benchmark, supervised fine-tuning (SFT) provides a strong baseline, substantially improving argument language consistency and end-to-end function call accuracy. Under consistent model selection, SFT achieves performance comparable to, and sometimes exceeding more complex reinforcement learning (RL) approaches. We further examine whether RL with structured, argument-aware rewards offers additional benefits. While methods such as Group Relative Policy Optimization (GRPO) can improve language consistency and better preserve general reasoning ability, these gains are incremental and most pronounced in generalization and multi-objective trade-offs. Overall, our results suggest that much of the performance in multilingual API grounding can be achieved through careful supervised training, with RL providing targeted rather than fundamental improvements.