AI 中文总结
研究在拉取请求打开时预测其接受情况及评审工作量的可行性,利用AIDev数据集构建预测管道,评估多种机器学习模型,发现接受度预测可行,评审工作量预测较难,早期模型可支持分类和优先级排序,应作咨询工具。
AI 中文摘要
拉取请求是现代软件仓库中审查和集成代码更改的核心机制。随着人工智能编码智能体开始与人类开发者一起提交更多代码更改,维护者面临新挑战:在评审讨论、持续集成反馈或合并决策之前,判断哪些拉取请求可能被接受,哪些可能需要大量评审工作。本文研究在拉取请求打开时能否估计这些结果。利用AIDev数据集,构建了一个针对人类和智能体编写的拉取请求的泄漏感知预测管道。特征集限于提交时信息。评估了多种经典机器学习模型,结果表明接受度预测从早期信号来看是可行的,基于树的模型F1分数高于0.95,文本清晰度和元数据是最有影响力的预测因素之一。评审工作量预测更难,提交时特征对评论计数和合并时间的解释有限,表明评审人员可用性、项目工作流程和团队特定评审实践起主要作用。这些发现表明早期拉取请求模型可支持分类和评审人员优先级排序,但应用作咨询工具而非自动决策者。
英文摘要
Pull requests (PRs) are a central mechanism for reviewing and integrating code changes in modern software repositories. As AI coding agents begin to submit more code changes alongside human developers, maintainers face a new challenge: deciding which PRs are likely to be accepted and which ones may require substantial review effort. This paper studies whether such outcomes can be estimated at the time a PR is opened, before reviewer discussion, CI feedback, or merge decisions are available. Using the AIDev dataset, we construct a leakage-aware prediction pipeline for human- and agent-authored PRs. The feature set is limited to submission-time information, including PR text characteristics, metadata, repository context, temporal signals, and lightweight diff statistics. We evaluate classical machine-learning models, including Logistic Regression, Random Forests, Gradient Boosting, Extra Trees, and MLPs, across pooled, human-only, agent-only, and balanced contributor views. Our results show that acceptance prediction is feasible from early signals: tree-based models achieve F1 scores above 0.95, with textual clarity and metadata among the most influential predictors. Review-effort prediction is more difficult. Comment counts and time-to-merge are only modestly explained by submission-time features, suggesting that reviewer availability, project workflow, and team-specific review practices play a major role. These findings indicate that early PR models can support triage and reviewer prioritization, but should be used as advisory tools rather than automated decision-makers.