发表机构
Virginia Tech; University of Georgia(弗吉尼亚理工大学; 佐治亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出控制-计算契约协议,解决智能体AI-RAN中LLM智能体与PHY处理共置时的双向耦合问题,通过三规则五阶段工作流将PHY时隙违规率从20.7%降至0.9%以下,并实现更快的gNB控制。
AI 中文摘要
人工智能无线接入网(AI-RAN)的最新进展正在将大型语言模型(LLM)驱动的智能体置于开放无线接入网(O-RAN)控制层级中。这种智能体AI-RAN的一个有前景的部署方案是将LLM推理与物理层(PHY)通信处理共同部署在相同的加速计算池上,以实现基础设施共享、数据本地化和低控制延迟。然而,这种共置会引发双向的控制-计算耦合,因为智能体与PHY处理竞争计算资源,同时其推理出的gNB控制动作可能改变未来的PHY工作负载,从而影响其下一次推理可用的计算资源。为此,本文提出了控制-计算契约,一种将AI-RAN工作负载治理与O-RAN控制执行相连接的协调协议。具体而言,我们首先在OTA O-RAN测试平台上揭示了直接的计算争用,其中持续的LLM推理使PHY解码时间增加了约八倍。然后,我们将该契约制定为三条具有运营商需求保护的协议规则,并将其映射到O-RAN和AI-RAN功能上,形成五阶段工作流。随后,一个基于OTA校准的模拟上行链路gNB控制案例研究表明,该契约将违反PHY时间预算的时隙比例保持在0.9%以下,而未经治理的共置情况下这一比例为20.7%,同时实现了比固定计算预留基准更快的gNB控制。最后,我们指出了其实际部署中的若干开放问题和展望。
英文摘要
Recent advances in artificial intelligence-radio access network (AI-RAN) are placing large language model (LLM)-driven agents within the open RAN (O-RAN) control hierarchy. A promising deployment for this agentic AI-RAN co-locates LLM inference with physical-layer (PHY) communication processing over the same accelerated compute pool for infrastructure sharing, data locality, and low control latency. However, this co-location induces bidirectional control-compute coupling, as the agent competes with PHY processing for compute, while its reasoned gNB control actions may alter future PHY workload and hence the compute available to its next inference. To this end, this article proposes the control-compute contract, a coordination protocol linking AI-RAN workload governance with O-RAN control execution. Specifically, we first expose direct compute contention on an over-the-air (OTA) O-RAN testbed, where continuous LLM inference increases the PHY decoding time approximately eightfold. We then formulate the contract as three protocol rules with operator-requirement protection, and map it onto O-RAN and AI-RAN functions as a five-stage workflow. Afterward, an OTA-calibrated case study of simulated uplink gNB control indicates that the contract keeps the proportion of slots violating the PHY time budget below $0.9\%$, against $20.7\%$ under ungoverned co-location, while achieving faster gNB control than a fixed compute reservation benchmark. Finally, we identify several open issues and outlooks for its practical deployment.
Comments7 pages, 4 figures. This work has been submitted to the IEEE for possible publication