arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00940cs.CL

用于建模迭代问题解决的数据集

A Dataset for Modeling Iterative Problem-Solving

  • Stanford University(斯坦福大学)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Fagun Patel, Sang T. Truong, Duc Q. Nguyen, Kazunori Fukuhara, Benjamin W. Domingue, Sanmi Koyejo, Nick Haber

AI总结:

该研究构建了包含3286名本科生超300万次代码提交的CodeInsight数据集,在此基准上评估了RSSM和LLM等模型,发现RSSM预测性能更优,LLM更适合作为生成式求解器。

AI中文摘要:

通过反复尝试解决问题是一项序列建模任务:每一步,求解者都会收到反馈并决定如何修改其解决方案。预测不同尝试间的性能是提升、停滞还是退化,是理解人类学习者和自主智能体的任何迭代问题解决过程的核心。除结果外,对哪些错误持续存在以及策略如何随尝试变化进行建模,能更深入地洞察序列学习的机制。研究这些动态需要观察大量求解者在尝试、接收反馈和修改的过程。具备自动评分功能的编程课程提供了这种场景,因为学生会向测试套件迭代提交代码并在每次尝试时接收反馈。因此,我们整理了CodeInsight,这是一个包含2个学年内2门C++入门课程的3286名本科生的超过300万次提交的大规模数据集,包含测试用例级别的结果、时间戳和源代码。在该数据集上,我们构建了一个基准,在统一的校准与评分协议下评估参数、序列和生成式传统的模型,包括适配了离散隐变量以跟踪求解者特征的循环状态空间模型(RSSM),以及生成显式解决方案的基于大语言模型(LLM)的预测器。适配后的RSSM在4门课程中的3门上实现了最强的预测准确率;LLM预测器准确率较低,但会在每次尝试时生成完整提交,支持直接分析失败模式。我们发现,在此场景中,模型的编码熟练度与预测性能呈负相关,LLM更适合被理解为基于上下文的生成式求解器,而非求解者行为的忠实预测器。我们应要求公开代码和数据集,以促进未来研究。

英文摘要:

Solving problems through repeated attempts is a sequential modeling task: at each step, the solver receives feedback and decides how to revise their solutions. Predicting whether performance improves, plateaus, or regresses across attempts is central to understanding any iterative problem-solving process in both human learners and autonomous agents. Beyond outcomes, modeling what errors persist and how strategies shift across attempts provides deeper insight into the mechanics of sequential learning. Studying these dynamics requires observing many solvers as they attempt, receive feedback, and revise. Programming courses with automated grading provide this setting, as students iteratively submit code to test suites and receive feedback on every attempt. We therefore curate CodeInsight, a large-scale dataset of over 3 million submissions from 3,286 undergraduates across 2 introductory C++ courses in 2 academic years, with test-case-level outcomes, timestamps, and source code. On this dataset, we build a benchmark that evaluates models spanning parametric, sequential, and generative traditions under a shared calibration-and-scoring protocol, including a Recurrent State Space Model (RSSM) adapted to track solver characteristics through discrete latent variables and an LLM-based predictor that generates explicit solutions. The adapted RSSM achieves the strongest predictive accuracy on three of the four courses. The LLM predictor is less accurate but produces full submissions at each attempt, enabling direct analysis of failure modes. We find that the model's coding proficiency is inversely related to predictive performance in this setting, with the LLM better understood as a generative solver conditioned on context rather than a faithful predictor of solver behavior. We publicly release our code and the dataset on request to facilitate future research.

补充信息

↑