arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05605cs.DC

性能可移植性研究中缺失数据的处理

Handling Missing Data in Performance Portability Studies

Ami Marowka

首次发表
浏览论文内容

中文总结 AI 辅助

针对性能可移植性研究中常见的缺失数据问题,本文评估多种插补方法,应用于真实HPC数据集,权衡准确性、鲁棒性与计算成本,为研究者提供实用建议。

中文摘要 AI 辅助

缺失数据是许多真实世界数据集中的常见挑战,往往导致结果偏差或精度降低。缺失数据在性能效率数据集中是一个普遍存在的问题,尤其是在高性能计算(HPC)的性能可移植性研究中。性能可移植性评分依赖于完整的性能效率集合作为其各自指标的输入;然而,由于基准测试不完整、硬件限制或实现缺口,可能会出现缺失值。许多已建立的性能可移植性指标缺乏处理缺失输入的机制,这可能导致分析不完整或完全无法计算评分。在本研究中,我们评估了一系列插补方法和算法,以解决缺失的性能效率问题。我们为每种方法提供了示例说明,并将其应用于真实的HPC性能可移植性数据集。我们的结果突出了准确性、鲁棒性和计算成本之间的权衡,为研究人员在性能可移植性评估中减轻缺失数据影响提供了实用建议。

英文摘要

Missing data is a common challenge in many real-world datasets, often leading to biased results or reduced accuracy. Missing data is a pervasive problem in performance efficiency datasets, particularly in high-performance computing (HPC) performance portability studies. Performance portability scores rely on complete sets of performance efficiencies as inputs to their respective metrics; however, missing values can arise due to incomplete benchmarking, hardware constraints, or implementation gaps. Many established performance portability metrics lack mechanisms for handling missing inputs, which can result in incomplete analyses or the inability to compute scores altogether. In this study, we evaluate a range of imputation methods and algorithms for addressing missing performance efficiencies. We present illustrative examples for each approach and apply them to real-world HPC performance portability datasets. Our results highlight the trade-offs between accuracy, robustness, and computational cost, offering practical recommendations for researchers seeking to mitigate the impact of missing data in performance portability assessments.

发表机构

  • Parallel Research Lab(平行研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑