arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19291cs.SE

EmbeddedKittens:对Scratch代码嵌入的评估

EmbeddedKittens: An Evaluation of Code Embeddings for Scratch

Benedikt Fein, Gordon Fraser

首次发表
浏览论文内容

中文总结 AI 辅助

研究探讨哪种代码嵌入方法适用于Scratch编程教育。实例化四个LLMs和五种嵌入方法,创建任务及数据集进行评估,结果表明在开放Scratch数据集上训练的嵌入模型可捕获代码信息用于学习分析,无需特定任务微调。

中文摘要 AI 辅助

将源代码嵌入用于机器学习应用的趋势也为编程教育中的学习分析带来了新机遇,但哪种代码嵌入方法最适合学习分析仍是个悬而未决的问题。常见的源代码嵌入方法是在训练大语言模型(LLMs)时将代码视为类似于自然语言的令牌序列。然而,对于像Scratch这样基于视觉块的编程语言,这种方法不能直接应用。本文实例化了四个LLMs和五种不同的Scratch程序嵌入方法,创建了令牌预测和两个不同的分类任务及相应数据集,并对模型进行了实证评估。实验表明将代码嵌入转移到Scratch教育环境是可行的。在大型开放Scratch数据集上训练的嵌入模型能够捕获有关代码的相关结构和语义信息,从而在典型的小课堂环境中实现学习分析,如预测学生程序的功能正确性,而无需进一步进行特定任务的模型微调。

英文摘要

The trend of embedding source code for machine learning applications also enables new opportunities in learning analytics in programming education, but which code embedding approach is most suitable for learning analytics remains an open question. A common approach to embedding source code lies in treating the code as a token sequence similar to natural language when training large language models~(LLMs). However, in case of visual block-based programming languages like Scratch, this approach cannot be applied directly. While text-based representations of block-based code can be created to apply LLMs to this problem, other dedicated embedding models could potentially exhibit improved performance by capturing additional structural information. In this paper, we therefore instantiate four LLMs and five different popular embedding approaches for Scratch programs, create a token-prediction and two different classification tasks with corresponding datasets, and empirically evaluate the models on them. Our experiments demonstrate that a transfer of code embeddings to the educational environment of Scratch is feasible. The embedding models trained on large open Scratch datasets capture relevant structural and semantic information about the code to enable learning analytics like predicting functional correctness of student programs, in the typically small classroom setting without requiring further task-specific model fine-tuning.

补充信息

↑