arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越回溯:利用大语言模型(LLMs)实现编程错误的自适应解释

Beyond the Traceback: Using LLMs for Adaptive Explanations of Programming Errors

Alexandru-Radu Moraru, Shreyan Biswas, Ujwal Gadiraju

arXiv 2608.20896首次发表:更新:

AI 中文总结

该研究通过103人众包实验,探究LLM生成的两类编程错误解释对调试的影响,发现LLM提升主观评价但未显著改善客观调试性能,提出自适应AI反馈系统需转向动态调整。

AI 中文摘要

编程错误消息对软件开发至关重要,但新手程序员仍难以理解。虽然大语言模型(LLMs)可以将这些错误改写为更清晰的解释,但可读性的提升是否能改善客观调试性能,以及解释风格应如何与程序员技能相匹配,目前尚不清楚。我们开展了一项包含103名参与者的多阶段众包研究,评估针对不同技能水平、由LLM生成的Python错误消息。通过自定义熟练度评估,我们将参与者按技能水平分类,测试了标准解释器消息与两种LLM生成风格:实用型(面向行动)和条件型(支架式解释)。我们测量了客观调试指标(修复率、尝试次数、修复时间)和主观感知(可读性、认知负荷、语气)。结果显示,LLM改写的消息显著提升了主观评价,其中实用型消息被评为更清晰且认知负荷更低,但这些感知上的提升并未转化为客观调试性能的统计显著改善。这凸显了一个关键的人机互补性差距:用户感觉更好的解释不一定能让他们成为更高效的调试者。我们讨论了自适应AI反馈系统的设计启示,认为未来工具应从静态的技能定向改写转向基于用户实时修复轨迹的动态调整。

英文摘要

Programming error messages are critical for software development, yet they remain difficult for novice programmers to interpret. While Large Language Models (LLMs) can rewrite these errors into clearer explanations, it remains unclear whether increased readability improves objective debugging performance or how explanation styles should align with programmer skill. We present a multi-stage crowdsourced study N=103 evaluating skill-targeted, LLM-generated Python error messages. Using a custom proficiency assessment, we categorized participants by skill level and tested standard interpreter messages against two LLM-generated styles: pragmatic (action-oriented) and contingent (scaffolded explanations). We measured both objective debugging metrics (fix rate, attempts, time-to-fix) and subjective perceptions (readability, cognitive load, tone). Our results show that while LLM-rewritten messages significantly improved subjective evaluations, with pragmatic messages rated as clearer and less cognitively demanding, these perceived gains did not translate into statistically significant improvements in objective debugging performance. This highlights a critical human-AI complementarity gap: explanations that feel better to users do not necessarily make them more effective debuggers. We discuss design implications for adaptive AI feedback systems, arguing that future tools should pivot from static skill-targeted rewriting toward dynamic adjustments based on a user's real-time repair trajectory.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑