arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PandasCorpus:真实世界Pandas工作流与使用模式资源

PandasCorpus: A Resource of Real-World Pandas Workflows and Usage Patterns

Syrym Abdikhan, Mazhar Hameed

arXiv 2608.14742首次发表:更新:

发表机构

Gisma University of Applied Sciences(吉斯马应用科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究构建了包含13.9万个Jupyter笔记本的PandasCorpus数据集,分析2015-2025年Pandas工作流特征,为相关研究提供公开资源。

AI 中文摘要

Pandas已成为数据处理和机器学习领域的事实标准库,广泛应用于数据加载、转换、分析等任务。尽管其应用十分普遍,但针对真实项目中Pandas的使用方式及实际典型工作流构成的系统研究却十分有限。为填补这一空白,我们推出PandasCorpus,这是一个从GitHub仓库整理而来的数据集,可大规模捕捉真实世界的Pandas工作流。本研究中,工作流指Jupyter笔记本中包含的基于Pandas的代码,Jupyter笔记本是编写、执行和共享数据分析代码的主流媒介。该数据集包含来自约10万个仓库的13.9万个笔记本,涵盖超过400万次Pandas API调用,涉及136种不同操作。除数据集构建外,我们还利用结构特征和Pandas特有的特征对工作流进行表征,并分析2015年至2025年间笔记本的演变情况。本研究考察了代码可执行性、笔记本大小以及Pandas操作的重复序列,为Pandas的实际使用提供了实证见解。生成的语料库为研究数据分析工作流、Pandas使用模式及感知库的代码组合提供了可重复使用的资源。该数据集和提取管道均可通过GitHub和Zenodo公开获取。

英文摘要

Pandas has emerged as the de facto library for data processing and machine learning, widely used for tasks, such as data loading, transformation, and analysis. Despite its ubiquity, there has been limited systematic investigation into how Pandas is used in real-world projects and how typical workflows are composed in practice. To address this gap, we introduce PandasCorpus, a dataset curated from GitHub repositories that captures real-world Pandas workflows at scale. In this work, a workflow refers to Pandas-based code contained in Jupyter notebooks, a prevalent medium for writing, executing, and sharing data analysis code. The dataset comprises 139k notebooks from approximately 100k repositories and captures more than 4M Pandas API calls spanning 136 distinct operations. Beyond dataset construction, we characterize workflows using structural and Pandas-specific features and analyze notebook evolution between 2015 and 2025. Our study examines code executability, notebook size, and recurring sequences of Pandas operations, providing empirical insights into how Pandas is used in practice. The resulting corpus offers a reusable resource for studying data analysis workflows, Pandas usage patterns, and library-aware code composition. Both the dataset and the extraction pipeline are publicly available via GitHub and Zenodo.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑