arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从讨论到执行:复制有缺陷和正确的数据科学代码

From Discussion to Execution: Replicating Buggy and Correct Data Science Code

Ragib Shahariar Ayon, Mohammad Wardat, Shibbir Ahmed

arXiv 2607.16569首次发表:更新:

AI 中文总结

研究旨在解决从问答论坛复制数据科学代码的难题,提出基于大语言模型的Reprodgen框架,能根据问题与答案生成可执行代码对,经多库评估及构建基准验证,实现可靠复制且模型性能有差异。

AI 中文摘要

从非正式来源复制可靠的数据科学代码具有挑战性,因为问题规范模糊、依赖项缺失和性能瓶颈。开发者问答论坛虽提供丰富讨论,但信息往往不完整且无结构,限制了其在自动调试和验证中的应用。本文介绍了Reprodgen,一个基于大语言模型的框架,用于从问答论坛帖子中自动复制可执行的有缺陷和已修复的数据科学程序。给定问题及其答案,Reprodgen重建问题中描述的错误行为和答案中描述的预期修复,生成反映原始讨论的可执行有缺陷和已修复代码对。该框架构建代码意图、功能需求和结构化思维链的结构化表示,并使用基于大语言模型的审查器迭代优化代码,直到可执行且语义一致。我们在包括pandas、numpy和scikit-learn在内的七个数据科学库的Stack Overflow和GitHub Issues上评估了Reprodgen,并构建了一个由人类专家验证的可运行有缺陷和已修复程序的基准。我们的管道使用大语言模型进行语义评估,同时通过实际执行验证可执行性。结果显示了可靠的复制,模型性能有明显差异。

英文摘要

Reproducing reliable data science code from informal sources is challenging due to ambiguous problem specifications, missing dependencies, and performance bottlenecks. Although developer Q&A forums provide rich discussions on diagnosing and fixing real-world issues, the information is often incomplete and unstructured, limiting its use for automated debugging and verification. In this paper, we introduce Reprodgen, a large language model (LLM) based framework for automatically replicating executable buggy and patched data science programs from Q&A forum posts. Given a question and its corresponding answer, Reprodgen reconstructs the buggy behavior described in the question and the intended fix described in the answer, producing executable buggy and patched code pairs that reflect the original discussion. The framework builds structured representations of code intent (CI), functional requirements (FR), and Structured Chain of Thought (SCoT), and iteratively refines code using an LLM-based reviewer until it is executable and semantically consistent. We evaluate Reprodgen on Stack Overflow (SO) and GitHub Issues (GI) across seven data science libraries, including pandas, numpy, and scikit-learn, and construct a benchmark of runnable buggy and patched programs validated by human experts. Our pipeline uses LLMs for semantic assessment, while executability is verified through actual execution. Results show reliable replication with clear differences in model performance.

Commentsaccepted at the 37th IEEE International Symposium on Software Reliability Engineering (ISSRE 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑