分布式计算连续体的性能指标分类法
A Taxonomy of Performance Metrics for the Distributed Computing Continuum
浏览论文内容
中文总结 AI 辅助
本文针对分布式计算连续体系统(DCCS)现有性能评估仅关注孤立维度的问题,提出结构化性能指标分类法,涵盖多维度指标及采集要求,为DCCS性能评估提供规范。
中文摘要 AI 辅助
性能评估对于理解、比较和改进计算系统至关重要,包括分布式计算连续体系统(Distributed Computing Continuum Systems,DCCS)。近年来,随着人工智能和大规模数据驱动应用的发展,计算需求发生了巨大变化,这些应用任务越来越多地分布在资源密集型数据中心和资源受限的边缘环境之间。在这种背景下,新型计算连续体架构和算法不断涌现,因此需要透明且一致的性能评估。然而,现有的评估实践往往只关注孤立的维度,如计算、网络、能效或应用级质量,因此仅能提供跨层DCCS行为的部分视图。本文提出了一种DCCS性能指标的结构化分类法,将指标分为计算级、网络级和应用/用户级类别,同时强调了可持续性、可观测性、适应性、数据局部性、迁移感知和连续体碎片化等新兴维度。此外,本文提供了数学公式,并讨论了它们与异构和动态连续体环境的相关性。本文还从采集范围、采集阶段和测量方法方面总结了指标采集要求,这些要求有助于明确指标是否可从单个节点、多个节点或整个系统采集,以及其更适合操作监控还是实验评估。
英文摘要
Performance evaluation is essential for understanding, comparing, and improving computing systems, including Distributed Computing Continuum Systems (DCCS). In recent years, computational requirements have changed substantially with the growth of artificial intelligence and large-scale data-driven applications. These application tasks are increasingly distributed between resource-intensive data centers and resource-constrained edge environments. In this context, novel computing continuum architectures and algorithms are emerging, creating a need for transparent and consistent performance evaluation. However, existing evaluation practices often focus on isolated dimensions, such as computation, networking, energy efficiency, or application-level quality, and therefore provide only a partial view of cross-layer DCCS behavior. This paper presents a structured taxonomy of performance metrics for DCCS. The taxonomy organizes metrics into computing-level, network-level, and application/user-level categories, while also highlighting emerging dimensions such as sustainability, observability, adaptability, data locality, migration awareness, and continuum fragmentation. Further, we provide mathematical formulations and discuss their relevance to heterogeneous and dynamic continuum environments. We also summarize metric acquisition requirements in terms of acquisition scope, acquisition phase, and measurement method. These requirements help clarify whether a metric can be collected from a single node, multiple nodes, or the full system, and whether it is more suitable for operational monitoring or experimental evaluation.