AI 中文总结
研究学生在编程教育中使用大语言模型辅导时,哪种回复助于高效学习。通过分析StudyChat数据集,用大语言模型辅助标注并经人工验证,发现回复风格与延续情况相关,支持情境感知的人工智能辅导回复评估和设计。
AI 中文摘要
随着学生在计算机科学教育中越来越多地使用大语言模型辅导,一个重要问题是:哪种回复有助于学生持续高效学习?以往研究了学生在计算机科学教育中如何使用大语言模型,但对辅导回复风格与编程求助情境中学生后续行为的关联了解较少。本文分析了StudyChat数据集,将其转化为203名学生的16851个助手-回复交互和2214次对话。通过本地大语言模型辅助注释进行标注,经人工验证与大语言模型辅助标签的一致性达82%。分析了整个数据集及不同求助情境下的有效延续和未解决延续情况。结果表明回复风格与有效延续和未解决延续显著相关,验证反馈的有效延续率最高,直接答案最低。还呈现了情境依赖的回复模式。这些发现支持了针对编程教育的人工智能辅导回复的情境感知评估和设计。
英文摘要
As students increasingly use LLM tutors in computer science education, one question becomes especially important: what kind of response helps a student continue productively? Prior work has studied how students use LLMs in computer science education, but less is known about how tutoring response styles are associated with student follow-up across programming help-seeking contexts. This paper analyzes StudyChat (UMass, 2026), a public dataset of student and ChatGPT tutoring conversations from an artificial intelligence course. We transformed StudyChat into 16,851 assistant-response interactions from 203 students and 2,214 conversations. Using local LLM-assisted annotation with Gemma 4, we labeled student help-seeking situations, student state, assistant response style, and student next-turn outcome. Human validation showed 82\% agreement with the LLM-assisted labels (Cohen's $κ=.74$). We analyzed productive continuation and unresolved continuation across the full dataset and across help-seeking contexts. Globally, response style was significantly associated with productive continuation, $χ^2(7)=100.39$, $p<.001$, $V=.078$, and unresolved continuation, $χ^2(7)=125.77$, $p<.001$, $V=.087$, though effect sizes were small. Verification feedback had the highest productive-continuation rate (82.4\%), while direct answers had the lowest (62.7\%). Descriptively, response-style score ranges were smallest in low-confusion conceptual contexts (.017) and largest in high-cognitive-load contexts (.203). More detailed comparisons showed situation-dependent response patterns. For example, stepwise guidance was followed by greater confusion decrease in high-cognitive-load code requests, while direct answers were followed by more unresolved continuation in high-load debugging. These findings support context-aware evaluation and design of AI tutoring responses for programming education.