arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33427cs.DC

迈向可组合云-HPC-边缘AI平台的系统之系统集成

Toward System-of-Systems Integration for Composable Cloud-HPC-Edge AI Platforms

Sumit Rakesh, Rajkumar Saini

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出将云-HPC-边缘AI平台视为系统之系统,通过基于接口、契约和策略的联邦化集成方法实现跨系统组合,并定义了边界测试、责任模型与七个集成面以支持互操作性和治理评估。

中文摘要 AI 辅助

现代AI平台日益整合基于不同假设构建的基础设施栈和运营模式,包括云风格服务平台、HPC工作负载管理系统、云原生编排、数据与工件系统、托管连接、可观测性以及边缘或信息物理环境。现有工作展示了选定栈之间的有效桥梁,但一种通用的推理跨独立控制系统组合的方式仍不成熟。我们认为,当独立有用的系统在贡献于更高级别AI平台能力的同时保留其自身的控制、管理、生命周期、策略和故障语义时,此类平台可被有效地视为系统之系统(SoS)。我们将可组合集成框定为一种基于接口、契约、映射、引用、策略上下文和运营证据的跨系统协调方法,同时保留原生控制平面并避免依赖单一拓扑或编排栈。由此产生的方向在使用上收敛,在控制上联邦化。本文提出了对等组成系统视图、用于区分组成系统与组件、局部依赖以及当前SoS边界之外独立有用系统的边界测试、局部、共享和限定责任模型、七个集成面以及一个代表性的跨系统工作流。最后,文章总结了用于评估互操作性、治理、可观测性、故障隔离、演进和复用的证据类别和研究问题。

英文摘要

Modern AI platforms increasingly combine infrastructure stacks and operating models designed around different assumptions, including cloud-style service platforms, HPC workload-management systems, cloud-native orchestration, data and artifact systems, managed connectivity, observability, and edge or cyber-physical environments. Existing work demonstrates effective bridges between selected stacks, but a general way to reason about composition across independently controlled systems remains underdeveloped. We argue that such platforms can be usefully viewed as systems of systems (SoS) when independently useful systems retain their own control, management, lifecycles, policies, and failure semantics while contributing to a higher-level AI platform capability. We frame composable integration as an approach to cross-system coordination based on interfaces, contracts, mappings, references, policy context, and operational evidence, while preserving native control planes and avoiding dependence on a single topology or orchestration stack. The resulting direction is converged in use and federated in control. The paper presents a peer constituent-system view, a boundary test for distinguishing constituent systems from components, local dependencies, and independently useful systems outside the current SoS boundary, a local, shared, and scoped responsibility model, seven integration surfaces, and a representative cross-system workflow. It concludes with evidence classes and research questions for evaluating interoperability, governance, observability, fault containment, evolution, and reuse.

发表机构

  • Luleå tekniska universitet(吕勒奥理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑