arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31647cs.LGcs.AI

开源游戏发布事故的仿真支持预测中的类型化时序交互特征

Typed Temporal Interaction Features for Simulation-Backed Forecasting of Open-Source Game Release Incidents

Shayma Alkobaisi, Anas Ali

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出类型化时序特征管道GAMEQUALGRAPH-Pilot,在仿真数据上预测开源游戏发布后30天内的质量事故,取得AUPRC 0.520,验证了协议的可复现性而非实际部署效果。

中文摘要 AI 辅助

开源视频游戏的质量取决于代码、资产、配置、测试、贡献者和问题工作流之间的交互,然而传统的缺陷预测器通常会扁平化或忽略这些关系。我们使用GAMEQUALGRAPH-Pilot研究了三十天内发布级质量事故的预测,这是一个类型化时序特征管道,具有校准的风险估计和努力感知排序。由于可访问的OS-SGameBench材料不提供人工审计的发布日期和爆发标签,所执行的评估明确是仿真支持的,而非关于真实游戏的经验性声明。五个种子世界各包含120个项目和24个发布,具有项目不相交的验证和未来的跨项目测试。该试点获得了0.520的AUPRC、0.673的AUROC、0.207的Brier分数,以及在百分之二十的测试预算下29.68%的努力感知召回率。其最接近的本地比较器Static-Hetero-Reimpl达到了0.522的AUPRC;在Holm校正后,-0.002的差异在统计上不显著。在测量环境中,每次发布的推理需要0.023毫秒。消融和受控缺失、漂移、引擎、项目规模、警报阈值和归因分析揭示了类型化交互在何处有帮助以及在何处失败。结果支持所提出协议的可复现性,而非部署有效性。在期刊提交或实际运营使用之前,真实的OSSGameBench发布重建、分层标签审计和官方图模型比较仍然是强制性的。这一边界保护了研究完整性并支持可信的评估。

英文摘要

Open-source video-game quality depends on inter-actions among code, assets, configuration, tests, contributors, and issue workflows, yet conventional defect predictors usually flatten or omit these relations. We investigate release-level forecasting of a quality incident within thirty days using GAMEQUALGRAPH-Pilot, a typed temporal feature pipeline with calibrated risk estimates and effort-aware ranking. Because the accessible OS-SGameBench materials do not provide manually audited release dates and outbreak labels, the executed evaluation is explicitly simulation-backed rather than an empirical claim about real games. Five seeded worlds each contain 120 projects and 24 releases, with project-disjoint validation and future cross-project testing. The pilot obtains an AUPRC of 0.520, AUROC of 0.673, Brier score of 0.207, and 29.68% effort-aware recall at a twenty-percent testing budget. Its closest local comparator, Static-Hetero-Reimpl, reaches 0.522 AUPRC; the -0.002 difference is not statistically significant after Holm correction. Inference requires 0.023 milliseconds per release in the measured environment. Ablations and controlled missingness, drift, engine, project-size, alert-threshold, and attribution analyses expose where typed interactions help and where they fail. Results support the reproducibility of the proposed protocol, not deployment effectiveness. Real OSSGameBench release reconstruction, stratified label audits, and official graph-model comparisons remain mandatory before journal submission or operational use in practice. This boundary protects research integrity and supports credible evaluation.

↑