arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MTAC-IFBench:多轮智能体编码中的指令遵循基准测试

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

Bosi Wen, Cunxiang Wang, Jiayi Gui, Haoke Zhang, Yilin Niu, Pei Ke, Dayong Yang, Hongning Wang, Minlie Huang

arXiv 2609.14992首次发表:更新:

发表机构

Tsinghua University; Zhipu AI; University of Electronic Science and Technology of China(清华大学; 智谱AI; 电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MTAC-IFBench基准,通过多轮渐进式开发指令和多样化约束评估LLM代码智能体的指令遵循能力,发现其性能随交互轮次增加而显著下降。

AI 中文摘要

近年来,大型语言模型(LLM)的快速发展通过使自主代码智能体能够迭代地规划、执行和利用外部工具来处理复杂任务,重塑了软件工程。除了实现功能正确性之外,这些智能体还必须在整个开发生命周期中忠实地遵循过程指令和约束。然而,现有基准通常侧重于最终功能正确性,或将指令遵循评估局限于单轮、通用对话或简单代码生成场景,使得多轮智能体编码中的指令遵循问题尚未得到充分探索。为弥补这一空白,我们提出了MTAC-IFBench,一个针对这一关键能力的全面基准。它包含多轮渐进式软件开发指令,涵盖6个主要类别和18个次要类别的多样化约束。每个实例平均有7.04轮和91.33个约束,对当前LLM构成了严峻挑战。为使评估可靠,我们为每个约束和功能需求构建了检查清单,并集成了验证脚本和评判智能体来验证每个检查清单项。MTAC-IFBench揭示了现有代码智能体在多轮指令遵循方面的显著缺陷,其性能随着交互会话的延长而迅速下降。

英文摘要

Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving functional correctness, these agents must faithfully follow process instructions and constraints throughout the development lifecycle. However, existing benchmarks typically focus on final functional correctness or confine instruction-following evaluation to single-turn, general chat or simple code generation scenarios, leaving instruction-following in multi-turn agentic coding underexplored. To bridge this gap, we propose MTAC-IFBench, a comprehensive benchmark for this critical capability. It features multi-turn progressive software development instructions with diverse constraints spanning 6 primary and 18 secondary categories. With an average of 7.04 turns and 91.33 constraints per instance, it poses a rigorous challenge to current LLMs. To make the evaluation reliable, we construct a checklist for each constraint and functional requirement, and integrate verification scripts and judge agents to verify each checklist item. MTAC-IFBench identifies significant deficiencies in existing code agents in multi-turn instruction-following, with their performance degrading rapidly as the interaction session grows longer.

Comments23 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑