arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Prolog学生编写的错误类型有哪些?实证分类法和数据驱动的变异框架

What Bugs Do Prolog Students Write? An Empirical Taxonomy and Data-Driven Mutation Framework

Ricardo Brancas, Pedro Orvalho, Carolina Carreira, Vasco Manquinho, Ruben Martins

arXiv 2607.21193首次发表:更新:

AI 中文总结

研究Prolog学生编写的错误类型,通过对7201份提交人工分类得出细粒度分类法,据此开发数据驱动的LogMorph变异工具,其操作符加权,评估显示合成错误分布与学生分布匹配,指出残留差异来源及改进方向。

AI 中文摘要

逻辑编程教育的自动化反馈工具依赖于反映学生实际错误的真实错误数据集。然而,现有的Prolog变异测试框架将所有变异视为同等可能,产生与课堂实际情况不同的合成错误。我们对265名本科生的7201份Prolog提交进行了实证研究,通过对200份错误修复提交进行人工分类,得出了学生错误的细粒度分类法。在此分类法的指导下,我们开发了LogMorph,一种数据驱动的变异工具,其17个操作符根据观察到的错误分布进行加权。LogMorph在抽象语法树上枚举有效的变异位点,按比例采样操作符,注入错误,在需要新代码片段时委托给基于SMT的合成器,并根据参考测试套件验证每个变异体。对16000个生成的变异体的评估表明,合成错误分布与学生分布密切匹配,大多数错误类别在两个百分点内一致。我们将与切割相关的变异和合成器生成的代码确定为残留差异的主要来源,并概述了将SMT后端与在学生代码上微调的语言模型相结合如何进一步提高真实性。

英文摘要

Automated feedback tools for logic programming education depend on realistic bug datasets that reflect the mistakes students actually make. However, existing mutation testing frameworks for Prolog treat all mutations as equally likely, producing synthetic faults that diverge from classroom reality. We present an empirical study of 7,201 Prolog submissions from 265 undergraduate students, from which we derive a fine-grained taxonomy of student bugs through manual classification of 200 bug-fixing submissions. Guided by this taxonomy, we develop LogMorph, a data-driven mutation tool whose 17 operators are weighted according to the observed error distribution. LogMorph enumerates valid mutation sites on the abstract syntax tree, samples operators proportionally, injects faults, delegating to an SMT-based synthesizer when new code fragments are needed, and validates each mutant against a reference test suite. An evaluation of 16,000 generated mutants shows that the synthetic error distribution closely matches the student distribution, with most bug categories agreeing to within two percentage points. We identify cut-related mutations and synthesizer-generated code as the main sources of residual divergence, and outline how combining the SMT back-end with a language model fine-tuned on student code can further improve realism.

CommentsIn Proceedings ICLP 2026, arXiv:2607.17707

Journal refEPTCS 450, 2026, pp. 163-178

DOI:10.4204/EPTCS.450.14

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑