arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当更好的回合并不造就更好的智能体:诊断下一回合指标与工作流成功之间的差距

When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success

Md Tahmid Rahman Laskar, Xue-Yong Fu, Gundeep Singh, Karol Chang, Kevin Sanders, Shi Zong, Tania Habib, Julien Bouvier Tremblay, Shayna Gardiner, Harsh Saini, Matthias Lee, Elena Khasanova, Quinten McNamara, Shashi Bhushan TN

arXiv 2609.21187首次发表:更新:

发表机构

Dialpad Inc.(Dialpad 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究诊断了下一回合指标与自主工作流成功之间的差距,发现SFT虽提升回合级性能,但无法转移至端到端工作流,最高成功率仅10.4%,故需分别报告各层面指标。

AI 中文摘要

智能体模型经常被逐决策评估,即模型基于黄金交互历史预测下一动作,并与参考动作进行评分。我们研究了这种协议下的改进是否能预测自主工作流执行的改进。我们研究了预SFT和监督微调(SFT)的Qwen3模型(4B和14B参数)以及Gemma 3模型(4B和12B参数)在多轮客户支持工作流上的表现。我们发现,SFT持续提高了文本回合成功率,并且在黄金历史评估下,每个模型的整体下一回合成功率都有所提升。然而,这些改进并未转移到自主工作流执行中。工具特定的收益也因指标和模型而异。四个SFT模型在整体工作流评估中均未成功,严格的轨迹完成率最高仅为10.4%的工作流成功率。我们的结果表明,下一回合评估不是工作流成功的可靠代理,这促使我们分别报告文本质量、局部动作正确性、工具执行和端到端任务完成情况。

英文摘要

Agent models are frequently evaluated one decision at a time, where the model predicts the next action based on the gold interaction history, which is scored against a reference. We investigate whether improvement under this protocol is predictive of improved autonomous workflow execution. We study pre-SFT and supervised fine-tuned (SFT) Qwen3 models at 4B and 14B parameters and Gemma 3 models at 4B and 12B parameters on multi-turn customer-support workflows. We find that SFT consistently improves text-turn success, and that overall next-turn success increases for every model under gold-history evaluation. However, these improvements do not transfer to autonomous workflow execution. Tool-specific gains also vary across metrics and models. None of the four SFT models succeeds under holistic workflow evaluation, with strict trajectory completion reaching at most 10.4% workflow success. Our results show that next-turn evaluation is not a reliable proxy for workflow success, motivating separate reporting of text quality, local action correctness, tool execution, and end-to-end task completion.

CommentsAccepted to the REALM Workshop at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑