计算网络科学中的可复现性挑战:证据、原因与建议
Reproducibility Challenges in Computational Network Science: Evidence, Causes, and Recommendations
- Leiden Institute of Advanced Computer Science Leiden University(莱顿大学高级计算机科学研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对计算网络科学中可复现性不足的问题,本文提出分类体系,通过四个案例研究揭示代码与数据缺失等成因,并给出强制共享工件、标准化基准等改进建议。
AI中文摘要:
可复现性对于科学进步至关重要,它使得验证、公平比较以及在先前工作基础上进行构建成为可能。然而,在计算网络科学(CNS)中,由于代码缺失、数据集不可访问以及实验细节报告不足,可复现性仍然有限。本文提出了一个CNS中可复现性的分类体系,该体系围绕工件可用性、算法清晰度、实验环境以及数据处理和实验流程进行构建。为了系统地审视这些挑战,我们开展了四项涵盖不同方法论设置的案例研究:基于主题的有影响力用户检测(网络科学与自然语言处理方法)、基于影响力的社区检测(纯网络方法)、影响力最大化(经典、启发式、近似及基于人工智能的混合方法),以及用于网络分析的强化学习(基于学习的方法)。在这些领域中,我们观察到公开可用工件(尤其是代码和数据集)的一致缺乏,这阻碍了结果的验证和比较。我们识别了导致这一可复现性差距的关键原因,包括共享工件的激励有限、数据访问限制、实验描述不完整以及复杂的方法论流程。最后,我们提出了改进可复现性的建议,包括强制性的工件共享政策、标准化基准以及实验设置的全面报告。解决这些差距对于确保计算网络科学的透明度、可比性和持续进步至关重要。
英文摘要:
Reproducibility is essential for scientific progress, enabling validation, fair comparison, and building upon prior work. In computational network science (CNS), however, reproducibility remains limited due to missing code, inaccessible datasets, and insufficient reporting of experimental details. This paper presents a taxonomy of reproducibility in CNS, structured around artifact availability, algorithmic clarity, experimental environments, and data processing and experimental pipelines. To systematically examine these challenges, we conduct four case studies spanning diverse methodological settings: topic-based influential user detection (network science and natural language processing-based methods), influence-based community detection (pure network-based methods), influence maximization (classical, heuristic, approximation, and AI-based mixed approaches), and reinforcement learning for network analysis (learning-based methods). Across these domains, we observe a consistent lack of publicly available artifacts, particularly code and datasets, hindering verification and comparison of results. We identify key causes of this reproducibility gap, including limited incentives for sharing artifacts, data access restrictions, incomplete experimental descriptions, and a complex methodological pipeline. Finally, we outline recommendations to improve reproducibility, including mandatory artifact sharing policies, standardized benchmarks, and comprehensive reporting of experimental setups. Addressing these gaps is critical to ensure transparency, comparability, and sustained progress in computational network science.