发表机构
LG CNS(LG CNS)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对韩国开放公共API的多步骤工具调用场景,构建KOPA-Bench基准测试,提出EDGE数据合成方案,微调9B模型后其性能接近同系列27B模型且在KOPA-Bench和BFCL上均有提升。
AI 中文摘要
数据主权法规日益要求公共机构部署开源本地部署的大语言模型(LLM)智能体,该智能体可跨实时政府API串联多个工具调用。然而,开源模型在这种多步骤场景中始终表现不佳,且现有基准测试未衡量该差距。我们推出韩国开放公共API基准测试(KOPA-Bench),包含145项现实任务。为缩小此差距,我们提出EDGE,即基于执行的动态图,用于由实时执行驱动的工具调用数据合成。EDGE构建每个工具输出如何作为另一个工具输入的图,仅保留针对实时API实际调用时成功的链接,并遍历这些已验证链接以合成可执行的多步骤轨迹。通过GRPO在所得数据集上进行微调后,我们的9B模型几乎与同系列未微调的27B模型相匹配,不仅在KOPA-Bench上大幅提升,在BFCL基准测试上也有显著改善。
英文摘要
Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise LLM agents that chain multiple tool-calls across live government APIs. However, open-source models consistently underperform in this multi-step setting, and no existing benchmark measures the gap. We introduce the Korean Open Public API Benchmark (KOPA-Bench), comprising 145 real-world tasks. To close this gap, we present EDGE, an Execution-grounded Dynamic Graph for tool-calling data synthEsis driven by live execution. EDGE builds a graph of how each tool's output can feed another's input, keeps only the links that succeed when actually called against the live APIs, and traverses these verified links to synthesize executable multi-step trajectories. Fine-tuned via GRPO on the resulting dataset, our 9B model nearly matches the untuned 27B model from the same family, improving substantially not only on KOPA-Bench but also on the BFCL benchmark.
Comments30 pages, 7 figures, 26 tables. Accepted to EMNLP 2026 Industry Track