发表机构
Jožef Stefan Institute(约瑟夫·斯特凡研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对边缘-云连续体中智能体AI部署问题,提出agentic-eCAL指标,结合能量模型与OSI传输,发现智能体间文本传输能耗仅占0.25%,主要成本来自通信引发的推理与上下文处理。
AI 中文摘要
随着电信网络向自主的5G-Advanced和6G运营演进,智能体人工智能(AI)工作流——其中大语言模型(LLMs)执行多步推理、调用诊断工具、检索领域知识并跨智能体团队协调——正日益嵌入边缘-云连续体中。尽管生物大脑以约20W的极低代谢功率预算完成复杂认知,当代LLM却极其耗能和耗内存,使得可持续的生命周期编排成为关键运营优先事项。然而,现有AI生命周期指标仅评估孤立的单模型推理,或完全忽视多智能体执行图。因此,网络运营商缺乏基础模型来确定分布式智能体通信是否产生有意义的能量成本,以及智能体团队应物理驻留在边缘-云层的何处。为填补这一空白,我们提出agentic-eCAL,将AI生命周期能量成本(eCAL)指标推广至有向多智能体工作流,通过耦合闭式双速率单次调用能量模型(计算受限的预填充和内存受限的解码)与7层OSI数据传输。基于NVIDIA A100和H100上数百种GPU基准配置、16个开放权重模型和8种编排拓扑,我们验证了该指标的组成部分并研究工作流放置的影响。我们的发现表明,跨5G RAN、城域和光链路的智能体间文本传输仅占工作流能量的0.25%。因此,在边缘-云智能体放置中,分布的主要能量成本通常不是智能体间文本本身的传输,而是由该通信引发的额外推理和上下文处理。
英文摘要
As telecommunication networks evolve toward autonomous 5G-Advanced and 6G operations, agentic artificial intelligence (AI) workflows, where large language models (LLMs) execute multi-step reasoning, invoke diagnostic tools, retrieve domain knowledge, and coordinate across agent teams, are increasingly embedded across the edge-cloud continuum. While the biological brain accomplishes complex cognition on an exceptionally modest metabolic power budget of approximately 20W contemporary LLMs are profoundly energy- and memory-intensive, making sustainable lifecycle orchestration a critical operational priority. However, existing AI lifecycle metrics evaluate only isolated, single-model inferences or overlook multi-agent execution graphs entirely. Consequently, network operators lack foundational models to determine whether distributed agent communication incurs meaningful energy costs and where across edge-cloud tiers agent teams should physically reside. To address this gap, we introduce agentic-eCAL, generalizing the Energy Cost of AI Lifecycle (eCAL) metric to directed multi-agent workflows by coupling a closed-form two-rate single-call energy model (compute-bound prefill and memory-bound decode) with 7-layer OSI data transport. Grounded in hundreds of GPU benchmark configurations on NVIDIA A100 and H100, 16 open-weight models and 8 orchestration topologies, we validate components of the metric and study workflow placement implications. Our findings demonstrate that inter-agent text transport incurs 0.25% of workflow energy across 5G RAN, metro, and optical links. Therefore in edge-cloud agent placement the dominant energy cost of distribution is often not the transmission of inter-agent text itself, but the additional inference and context processing induced by that communication.