发表机构
School of Cyber Science and Engineering, Huazhong University of Science and Technology; Key Laboratory of Cyberspace Security, Ministry of Education; Hubei Key Laboratory of Distributed System Security(华中科技大学网络空间科学与工程学院; 教育部网络空间安全重点实验室; 湖北省分布式系统安全重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现有LLM竞赛编程基准的不足,推出含322道题的InteractBench基准,评估LLM在未公开信息交互式竞赛编程任务的表现,发现先进模型存在显著交互差距并提出失败分类法。
AI 中文摘要
竞赛编程正被越来越多地用于评估大语言模型(LLM)的算法推理能力,但现有基准主要聚焦于全信息任务——即所有问题输入预先提供,这忽略了算法推理的一个关键维度:生成的程序在关键信息未预先公开时的运行能力。竞赛编程中的交互式问题是体现这一挑战的典型部分,这类问题要求程序在严格协议约束和有限查询预算下,与交互器(评判程序)进行多轮交互,仅在查询后才会揭示新信息。为填补这一空白,我们推出InteractBench,该基准包含322道从Codeforces、AtCoder、IOI和ICPC中精选的高质量交互式问题,每个问题都配有可执行的本地交互器,支持完全离线评估。与现有基准不同,InteractBench评估模型生成的代码能否动态获取信息并跟踪状态。我们的评估揭示了显著的交互差距:即使是最先进的推理模型在交互式问题上的成功率也有限。除成功率外,我们还提出了细粒度失败分类法,以诊断这些缺陷的根本原因。尽管算法逻辑错误仍是主要问题,但协议违规和查询预算超支也频繁出现。代码可在this https URL获取。
英文摘要
Competitive programming is increasingly being used to evaluate the algorithmic reasoning capabilities of large language models (LLMs). However, existing benchmarks primarily focus on full-information tasks where all problem inputs are provided upfront. This overlooks a critical dimension of algorithmic reasoning: the ability of generated programs to operate when key information is not revealed upfront. Interactive problems, a distinctive component of competitive programming, embody this challenge. These problems require programs to engage in multi-round interaction with an interactor (a judge program) under strict protocol constraints and limited query budgets, with new information revealed only in response to queries. To address this gap, we introduce InteractBench, a benchmark comprising 322 high-quality interactive problems curated from Codeforces, AtCoder, IOI, and ICPC. Each problem is packaged with executable local interactors, enabling fully offline evaluation. Unlike existing benchmarks, InteractBench assesses whether model-generated code can acquire information and track state dynamically. Our evaluation reveals a significant interaction gap: even the most advanced reasoning models achieve limited success on interactive problems. Beyond success rates, we propose a fine-grained failure taxonomy to diagnose the root causes of these deficiencies. Although algorithmic logic errors remain dominant, protocol violations and query-budget overruns are frequent. Code is available at https://github.com/kmsgk0/InteractBench.
CommentsAccepted at ICML 2026