AI 中文总结
研究利用GitHub问题和拉取请求构建数据集时,传统流程证据分散难审计。RepoTrace是浏览器辅助工具,结合扩展、后端和仪表盘收集证据,验证表明可支持完整本地证据收集工作流程。
AI 中文摘要
实证软件工程研究常从GitHub问题和拉取请求构建数据集。许多项目中,研究人员在浏览器检查页面、复制字段、记录笔记,之后运行脚本处理数据。此工作流程灵活,但页面证据、研究代码和决策依据分散,难以审计。RepoTrace是浏览器辅助研究工具,将GitHub问题和拉取请求证据收集到本地SQLite支持的工作区。它结合Chrome侧边栏扩展、Express后端和React仪表盘来捕获页面快照、评论、标签、笔记、筛选和标记决策、刷新历史记录和范围导出,将源证据和研究解释联系在一起。在两个研究项目中对20个Matplotlib问题进行了验证。结果数据集保留了22个快照、38条评论、20条研究笔记、98条注释、20次筛选评论、20条修复证据条目和4个模拟未解决的共识冲突。结果表明,RepoTrace可以支持为手动构建的GitHub问题和拉取请求数据集提供完整的本地证据收集工作流程。
英文摘要
Empirical software engineering studies frequently build datasets from GitHub issues and pull requests. In many projects, researchers inspect pages in a browser, copy selected fields into spreadsheets, keep side notes in separate documents, and later run scripts to normalize or export the data. This workflow is flexible, but the page evidence, the research codes, and the rationale behind each decision end up spread across tabs and files, which leaves provenance, update tracking, and multi-reviewer labeling hard to audit. RepoTrace is a browser-assisted research tool that collects GitHub issue and pull-request evidence into a local SQLite-backed workspace. It combines a Chrome side-panel extension, an Express backend, and a React dashboard to capture page snapshots, comments, labels, notes, screening and labeling decisions, refresh history, and scoped exports, keeping the source evidence and the research interpretation linked together. A validation pass collected and checked 20 Matplotlib issues across two study projects. The resulting dataset preserves 22 snapshots, 38 comments, 20 research notes, 98 annotations, 20 screening reviews, 20 fix-evidence entries, and 4 simulated unresolved consensus conflicts. The results show that RepoTrace can support a complete local evidence-collection workflow for manually constructed GitHub issue and pull-request datasets.
Comments6 pages. Accepted to the ISSTA 2026 Tool Demonstrations Track; published in the Companion Proceedings of SPLASH Companion '26