arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17111cs.CYcs.AIcs.LO

使用LLM发现数学形式化建模中的常见错误

Finding Common Mistakes In Modelling With Mathematical Formalisms Using LLMs

Lilian Killich, Marko Schmellenkamp, Fabian Vehlken, Thomas Zeume

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出一种基于LLM的工作流程,用于识别数学形式化建模中的常见错误,通过错误修复转换生成候选并聚类可视化,在命题逻辑等数据集上验证了其有效性和可扩展性。

中文摘要 AI 辅助

使用逻辑公式、数学方程或正则表达式等数学形式化进行建模,对于计算机科学及其他STEM学科的学生来说,是一项重要但具有挑战性的任务。识别在此过程中出现的常见错误,是通过提供有针对性的高质量反馈(例如在交互式学习系统中)帮助困难学生的重要一步。我们提出了一种支持工具的工作流程,该流程能够:(1)识别可作为常见错误候选的、能够解释大型教育数据集中许多学生错误的模式,(2)根据相似性对候选模式进行聚类,以及(3)为教师和计算机教育研究人员可视化生成的聚类。该可视化旨在帮助研究人员识别常见的建模错误。常见错误的候选模式通过错误修复转换来表示,这些转换将不正确的形式化转换为正确的形式化;它们由LLM生成并通过算法验证。我们通过复现文献中手工识别的命题逻辑建模中的常见错误,展示了该方法效果良好;同时表明,与其他算法方法不同,基于LLM的方法适用于非常大的数据集;并将其应用于多种其他形式化,以展示其泛化能力超越命题逻辑。

英文摘要

Modelling with mathematical formalisms like logical formulas, mathematical equations, or regular expressions is an important yet challenging task for students of computer science and other STEM disciplines. Identifying common mistakes occurring in this context is an important step towards helping struggling students by providing targeted high-quality feedback, e.g. in interactive learning systems. We present a tool-supported workflow that allows to (1) identify candidates for common mistakes that explain many student mistakes in large educational data sets, (2) cluster candidates according to similarities, and (3) visualize resulting clusters for instructors and CS education researchers. The visualization is designed to help researchers to identify common modelling mistakes. The candidates for common mistakes are represented by bug fixing transformations that translate incorrect formalizations into correct formalizations; they are generated by an LLM and validated algorithmically. We show that this approach works well by reproducing common mistakes in propositional logic modelling that were identified by hand in the literature; showing that, unlike other algorithmic approaches, the LLM-based approach is suitable for very large sets of data; and applying it to multiple other formalisms to showcase it generalizes beyond propositional logic.

发表机构

  • Ruhr University Bochum(波鸿鲁尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑